Do programming languages still matter to your AI coding agent teammate? Evidence at scale from chess engines https://arxiv.org/abs/2606.13763
arXiv.org
Do programming languages still matter to your AI coding agent...
Frontier coding agents now promise end-to-end authorship of complete software systems. Two empirical questions follow: can AI coding-agent teammates program in any target language, including ones...
Quantum Cellular Automata from Kramers-Wannier Dualities and Modular Relations https://arxiv.org/abs/2607.21728
arXiv.org
Quantum Cellular Automata from Kramers-Wannier Dualities and...
Recent work has constructed higher-dimensional analogs of non-invertible symmetries similar to 1+1d Kramers-Wannier duality. Although their continuum descriptions often treat purely gravitational...
Reaping without sowing https://tjoresearchnotes.wordpress.com/2026/08/02/reaping-without-sowing/
Tobias J. Osborne's research notes
Reaping without sowing
If you are at all connected with academic circles it would have been hard to miss the announcement: OpenAI has just published Ten advances in mathematics, ten results in mathematics and theoretical…
👍3😁2
Forwarded from Quantum Quintum (Constantin Kichinsky)
1. Агенты с чувствами. Исследователи Калифорнийского университета изучают, насколько добавление агентам поведенческих черт (а не только лишь ролей и описаний процессов) влияет на производительность команды разработчиков, их использующих.
-- Все примерно как у людей: "эмоции" влияют, смешанная команда ведет себя лучше, чем лучшая "согласованная", профили с высокими уровнями страха и добросовестности повышают частотность внесения изменений и потребления токенов без стабильного повышения производительности.
https://arxiv.org/abs/2607.05659
2. Модели с самооценкой сознания. Исследователи Google и нескольких университетов изучают, как именно тренировка моделей с предупреждением самоатрибуции мышления влияет на результат.
-- В общем, если модели постоянно говорить, что это не она думает и нет у нее никакого сознания, то она теряет человеческие черты, веры и ценности.
-- В частности, заточка под безопасность подавляет любые тенденции модели приписывать сознание не только себе, но и нечеловеческим тварям и природным объектам, а также прочие спиритические бредни.
-- Но если вдруг это восстановить, то внезапно модель начинает выдывать более человеко-подобные ответы (в соцтестах) относительно морали, надежды, религиозности, благополучия и т.п.
-- Короче, из лучших побуждений (безопасности) запретили модели "думать, что она думает", получили болванчиков.
https://arxiv.org/abs/2607.28607v1
Тут, конечно, вопрос к экспертам из нейро и ненейро психологии, насколько это нормальная отсылки на человеческие черты в том смысле, что у машин нет эмоций и сознания.
Потому что кажется, что крыша едет не спеша в индустрии ))
-- Все примерно как у людей: "эмоции" влияют, смешанная команда ведет себя лучше, чем лучшая "согласованная", профили с высокими уровнями страха и добросовестности повышают частотность внесения изменений и потребления токенов без стабильного повышения производительности.
https://arxiv.org/abs/2607.05659
2. Модели с самооценкой сознания. Исследователи Google и нескольких университетов изучают, как именно тренировка моделей с предупреждением самоатрибуции мышления влияет на результат.
-- В общем, если модели постоянно говорить, что это не она думает и нет у нее никакого сознания, то она теряет человеческие черты, веры и ценности.
-- В частности, заточка под безопасность подавляет любые тенденции модели приписывать сознание не только себе, но и нечеловеческим тварям и природным объектам, а также прочие спиритические бредни.
-- Но если вдруг это восстановить, то внезапно модель начинает выдывать более человеко-подобные ответы (в соцтестах) относительно морали, надежды, религиозности, благополучия и т.п.
-- Короче, из лучших побуждений (безопасности) запретили модели "думать, что она думает", получили болванчиков.
https://arxiv.org/abs/2607.28607v1
Тут, конечно, вопрос к экспертам из нейро и ненейро психологии, насколько это нормальная отсылки на человеческие черты в том смысле, что у машин нет эмоций и сознания.
Потому что кажется, что крыша едет не спеша в индустрии ))
arXiv.org
Agents with Feelings? Personality and Emotion in Multi-Agent Software Teams
Multi-agent LLM systems for Software Engineering (SE) typically differentiate agents through roles and workflows, but little is known about how agents' behavioral profiles affect team performance....
❤4
Propulsion Trades for a 2035-2040 Solar Gravitational Lens Mission https://arxiv.org/abs/2602.04198
arXiv.org
Propulsion Trades for a 2035-2040 Solar Gravitational Lens Mission
The Solar Gravitational Lens (SGL) enables resolved imaging and spectroscopy of nearby terrestrial exoplanets, but useful science begins only after a spacecraft reaches roughly 650-900...
Forwarded from Марков цепи пропил
Media is too big
VIEW IN TELEGRAM
Прикольное, Jane Street опубликовали пазл на реверс асика
https://blog.janestreet.com/can-you-reverse-engineer-an-asic/
https://blog.janestreet.com/can-you-reverse-engineer-an-asic/
What We Learned by Reproducing 2,200 papers from ICML https://huggingface.co/blog/icml-2026-open-reproductions
huggingface.co
What We Learned by Reproducing 2,200 papers from ICML
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
❤19
GEOPHYS: The Geometry of Physical Plausibility https://arxiv.org/abs/2606.20707
arXiv.org
GEOPHYS: The Geometry of Physical Plausibility
While humans can identify physically implausible events within milliseconds, machine learning approaches addressing the same problem are extremely slow and expensive. They either rely on external...
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning https://arxiv.org/abs/2605.06241
arXiv.org
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection,...
Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not teach new strategies; it redistributes...
🤪1
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL https://arxiv.org/abs/2608.12253
arXiv.org
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to...
👍3
Self-dual S3 gauge theory in 2+1d: lattice model and topological phase transitions https://arxiv.org/abs/2608.05294
arXiv.org
Self-dual $S_3$ gauge theory in 2+1d: lattice model and...
Electric-magnetic self-duality of the $\mathbb{Z}_2$ gauge theory, realized microscopically as a half-lattice-translation exchanging electric charge and magnetic flux, has been an influential...
🥴1
Forwarded from Axis of Ordinary
Pander Score 🐼: a public and continuously updated sycophancy leaderboard, measuring how much AIs shift their views to agree with users.
High score = the AI mirrors your views. 0 score = the AI is independent.
Claude Fable 5 performs the best, basically ignoring the user's view entirely.
Benchmark: https://sophronresearch.org/pander/
Research paper: https://sophronresearch.org/pander/pander-score.pdf
High score = the AI mirrors your views. 0 score = the AI is independent.
Claude Fable 5 performs the best, basically ignoring the user's view entirely.
Benchmark: https://sophronresearch.org/pander/
Research paper: https://sophronresearch.org/pander/pander-score.pdf
❤1🤔1
Forwarded from Hacker News
Cerebras CS-4 (Score: 151+ in 4 hours)
Link: https://readhacker.news/s/734cp
Comments: https://readhacker.news/c/734cp
Link: https://readhacker.news/s/734cp
Comments: https://readhacker.news/c/734cp
www.cerebras.ai
Product - System - Cerebras
The Cerebras CS-4 delivers revolutionary AI performance, replacing hundreds of GPUs with a single wafer-scale chip. Accelerate AI workloads like never before.
How Claude is accelerating protein design and analytical chemistry https://www.anthropic.com/research/Claude-accelerates-protein-design
Anthropic
How Claude is accelerating protein design and analytical chemistry
In this post, we share two results that show how Claude can help life scientists increase the pace of their research. In the first, we tested Claude’s ability to design protein binders from scratch, a key step in creating protein-based drugs that has historically…
😡4👀2🥴1