OpenAI introduced GPT-Rosalind, a frontier reasoning model built to support research across biology, drug discovery, and translational medicine.
GPT-Rosalind is optimized for scientific workflows, with stronger performance in protein and chemical reasoning, genomics analysis, biochemistry knowledge, and scientific tool use.
GPT-Rosalind is optimized for scientific workflows, with stronger performance in protein and chemical reasoning, genomics analysis, biochemistry knowledge, and scientific tool use.
OpenAI
Introducing GPT-Rosalind for life sciences research
OpenAI introduces GPT-Rosalind, a frontier reasoning model built to accelerate drug discovery, genomics analysis, protein reasoning, and scientific research workflows.
💯2🔥1👏1💊1
Cool work by Google. Team built an AI system that discovers health biomarkers from wearable data: CoDaS
One of its first findings: "late-night doomscrolling" is a statistically validated predictor of depression severity (ρ = 0.177, p < 0.001, n = 7,497).
The AI named the feature. No human guidance.
CoDaS is a multi-agent system that runs the full biomarker discovery lifecycle autonomously:
Sensor data → Generate hypotheses → Run statistical + ML analysis → Conduct adversarial validation → Write manuscript
Research team deploy it across 9,279 participants and 3 clinical cohorts.
Here's where it gets interesting.
On one cohort, CoDaS found a feature with R² = 0.963. Near-perfect prediction of insulin resistance, passing 10/11 tests.
Then the AI rejected it. Finding it was glucose², a tautological transform of the target. True R² after removal: 0.389.
Researchers ran a blind expert evaluation. 15 domain experts, 76 manuscript assessments.
CoDaS: 86% acceptance rate, AI Co-Scientist: 85% rejection rate, Data Science Agent: 95% rejection rate,
Biomni: 100% rejection rate
No baseline received a single Accept or Minor Revision.
A surprising result: CoDaS found circadian instability features in two separate depression cohorts.
Sleep duration variability in one (ρ = 0.252). Sleep onset variability in the other (ρ = 0.126).
The cohorts were processed completely independently.
CoDaS compressed ~37 person-days of research (expert estimate) into 6-8 hours.
But the point isn't speed. It's that separating exploration from adversarial validation at the architecture level produces biomarker candidates that domain experts rate as scientifically valid.
One of its first findings: "late-night doomscrolling" is a statistically validated predictor of depression severity (ρ = 0.177, p < 0.001, n = 7,497).
The AI named the feature. No human guidance.
CoDaS is a multi-agent system that runs the full biomarker discovery lifecycle autonomously:
Sensor data → Generate hypotheses → Run statistical + ML analysis → Conduct adversarial validation → Write manuscript
Research team deploy it across 9,279 participants and 3 clinical cohorts.
Here's where it gets interesting.
On one cohort, CoDaS found a feature with R² = 0.963. Near-perfect prediction of insulin resistance, passing 10/11 tests.
Then the AI rejected it. Finding it was glucose², a tautological transform of the target. True R² after removal: 0.389.
Researchers ran a blind expert evaluation. 15 domain experts, 76 manuscript assessments.
CoDaS: 86% acceptance rate, AI Co-Scientist: 85% rejection rate, Data Science Agent: 95% rejection rate,
Biomni: 100% rejection rate
No baseline received a single Accept or Minor Revision.
A surprising result: CoDaS found circadian instability features in two separate depression cohorts.
Sleep duration variability in one (ρ = 0.252). Sleep onset variability in the other (ρ = 0.126).
The cohorts were processed completely independently.
CoDaS compressed ~37 person-days of research (expert estimate) into 6-8 hours.
But the point isn't speed. It's that separating exploration from adversarial validation at the architecture level produces biomarker candidates that domain experts rate as scientifically valid.
❤4💯3🥰2
New from DeepSeek: Mega MoE
Instead of running MoE as a chain of separate steps (dispatch → MLP → combine), Mega MoE fuses everything into a single mega-kernel. Even more importantly, it overlaps NVLink communication with Tensor Core computation, reducing the classic “compute–wait–transfer” bottleneck.
The result is a shift from fragmented execution to a continuous pipeline: higher GPU utilization, less idle time, and much better scaling in multi-GPU MoE workloads.
What’s also interesting is the direction: alongside this, DeepSeek is exposing low-level controls (SM usage, Tensor Core utilization, JIT behavior), turning DeepGEMM into a tunable performance toolkit, not just a fast library.
Feels less like a feature drop, more like a rewrite of how MoE is executed at scale.
Instead of running MoE as a chain of separate steps (dispatch → MLP → combine), Mega MoE fuses everything into a single mega-kernel. Even more importantly, it overlaps NVLink communication with Tensor Core computation, reducing the classic “compute–wait–transfer” bottleneck.
The result is a shift from fragmented execution to a continuous pipeline: higher GPU utilization, less idle time, and much better scaling in multi-GPU MoE workloads.
What’s also interesting is the direction: alongside this, DeepSeek is exposing low-level controls (SM usage, Tensor Core utilization, JIT behavior), turning DeepGEMM into a tunable performance toolkit, not just a fast library.
Feels less like a feature drop, more like a rewrite of how MoE is executed at scale.
GitHub
[Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes by LyricZhao · Pull Request #304 · deepseek-ai/DeepGEMM
New features
Mega MoE, fusing & overlapping dispatch/linear 1/SwiGLU/linear 2/combine into a single mega-kernel, overlapping NVLink communication and tensor core computation
Performance n...
Mega MoE, fusing & overlapping dispatch/linear 1/SwiGLU/linear 2/combine into a single mega-kernel, overlapping NVLink communication and tensor core computation
Performance n...
❤5👍3💯2
Redwood Research presents LinuxArena
- 20 live production environments for AI agents
- Frontier models achieve ~23% undetected sabotage vs. trusted monitors
- Useful work ≈ attack surface → sandboxing fails, monitoring is essential
- 20 live production environments for AI agents
- Frontier models achieve ~23% undetected sabotage vs. trusted monitors
- Useful work ≈ attack surface → sandboxing fails, monitoring is essential
🔥2👏1💯1
Sakana AI discovered a prompting technique named String Seed of Thought (SSoT).
This work will be presented at ICLR2026
The method is simple: instruct the LLM to generate a random string in its own output, then manipulate that string to derive its answer.
It requires only a small addition to the prompt and no external random number generator a prompting technique named String Seed of Thought (SSoT).
SSoT significantly reduces output bias across a wide range of LLMs, both open and closed.
With reasoning models (such as DeepSeek-R1), it reaches accuracy close to that of actual random sampling.
The method generalizes from binary choices to n-way selections and arbitrary probability distributions.
On the NoveltyBench diversity benchmark, SSoT outperformed other approaches across all six categories while maintaining output quality.
This work will be presented at ICLR2026
The method is simple: instruct the LLM to generate a random string in its own output, then manipulate that string to derive its answer.
It requires only a small addition to the prompt and no external random number generator a prompting technique named String Seed of Thought (SSoT).
SSoT significantly reduces output bias across a wide range of LLMs, both open and closed.
With reasoning models (such as DeepSeek-R1), it reaches accuracy close to that of actual random sampling.
The method generalizes from binary choices to n-way selections and arbitrary probability distributions.
On the NoveltyBench diversity benchmark, SSoT outperformed other approaches across all six categories while maintaining output quality.
Sakana AI
String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
SSoT: A simple prompting method that enables LLMs to generate distribution-faithful and diverse outputs.
🙏3💯2🥰1
AI4Science Catalyst backed by researchers from Stanford and Princeton unveiled LabWorld Factory, a "world engine" for biology.
The platform allows developers to generate fully scalable, simulated 3D biology labs entirely from natural language prompts.
Starting with a base of over 100 lab assets, the engine procedurally generates diverse layouts, physical tools, reagents, and camera views.
The platform allows developers to generate fully scalable, simulated 3D biology labs entirely from natural language prompts.
Starting with a base of over 100 lab assets, the engine procedurally generates diverse layouts, physical tools, reagents, and camera views.
❤3
Anthropic: Conway will evolve always on agents to the next level
Imagine an always-on Agent with custom UI tabs that users can share and reuse as packages. Mission control, any custom workflow that requires a UI, etc.
And all these to be powered by top models from Antropic. This is what "Claude Conway" will likely be about.
> Anthropic continues working on its always-on agent, Conway, with a new setting UI being added to the iOS app (currently hidden).
> On the web, a new UI component for Built-in and Installed has been introduced.
> Since we know new extensions will allow users to build custom UI tabs, we might be talking about a huge new feature here.
It is cooking😜
Imagine an always-on Agent with custom UI tabs that users can share and reuse as packages. Mission control, any custom workflow that requires a UI, etc.
And all these to be powered by top models from Antropic. This is what "Claude Conway" will likely be about.
> Anthropic continues working on its always-on agent, Conway, with a new setting UI being added to the iOS app (currently hidden).
> On the web, a new UI component for Built-in and Installed has been introduced.
> Since we know new extensions will allow users to build custom UI tabs, we might be talking about a huge new feature here.
It is cooking
Please open Telegram to view this post
VIEW IN TELEGRAM
TestingCatalog
Anthropics works on its always-on agent with new UI extensions
Anthropic is developing Conway, an always-on Claude agent with a containerized setup and full parity in settings, coming to both web and mobile.
Cloudflare open-sourced an email client where an AI agent reads your inbox, drafts your replies, and never sends anything without your permission.
It's called Agentic Inbox. It runs entirely on Cloudflare Workers. Zero third-party servers touching your emails.
It's called Agentic Inbox. It runs entirely on Cloudflare Workers. Zero third-party servers touching your emails.
GitHub
GitHub - cloudflare/agentic-inbox: A self-hosted email client with an AI agent, running entirely on Cloudflare Workers
A self-hosted email client with an AI agent, running entirely on Cloudflare Workers - cloudflare/agentic-inbox
🆒3
HuggingFace introduced ml-intern, the agent that just automated the post-training team
It's an open-source implementation of the real research loop that ML researchers do every day.
You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem.
How it works?
ml-intern makes full use of the HF ecosystem:
- finds papers on arxiv and hf.co/papers, reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on hf.co/datasets
- browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data
- launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains
ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like.
CLI
Web + mobile
Also provisioned 1k$ GPU resources and Anthropic credits for you to use.
It's an open-source implementation of the real research loop that ML researchers do every day.
You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem.
How it works?
ml-intern makes full use of the HF ecosystem:
- finds papers on arxiv and hf.co/papers, reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on hf.co/datasets
- browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data
- launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains
ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like.
CLI
Web + mobile
Also provisioned 1k$ GPU resources and Anthropic credits for you to use.
huggingface.co
Daily Papers - Hugging Face
Your daily dose of AI research from AK
🆒5
This is cool. Meet Simula from Google
A Google framework produces better synthetic datasets.
Instead of just randomly generating data or copying existing examples, Simula uses AI to strategically architect an entire dataset from scratch.
A Google framework produces better synthetic datasets.
Instead of just randomly generating data or copying existing examples, Simula uses AI to strategically architect an entire dataset from scratch.
Google Research
Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles
To address the scarcity of data required for specialized AI, we introduce Simula, a framework that reframes synthetic data generation as dataset-level mechanism design. By using reasoning to architect datasets from first principles, Simula enables fine-grained…
❤4
Anthropic's Mythos AI model is being accessed by unauthorized users 🤣
A handful of users in a private online forum gained access to Mythos on the same day that Anthropic first announced a plan to release the model to a limited number of companies for testing purposes, said the person, who asked not to be named for fear of reprisal. The group has been using Mythos regularly since then, though not for cybersecurity purposes, said the person, who corroborated the account with screenshots and a live demonstration of the model.
The users relied on a mix of tactics to get into Mythos. These included using access the person had as a worker at a third-party contractor for Anthropic and trying commonly used internet sleuthing tools often employed by cybersecurity researchers, the person said. The users are part of a private Discord channel that focuses on hunting for information about unreleased models, including by using bots to scour for details that Anthropic and others have posted on unsecured websites such as GitHub.
“We’re investigating a report claiming unauthorized access to Claude Mythos Preview through one of our third-party vendor environments,” a spokesperson for Anthropic said in a statement.
A handful of users in a private online forum gained access to Mythos on the same day that Anthropic first announced a plan to release the model to a limited number of companies for testing purposes, said the person, who asked not to be named for fear of reprisal. The group has been using Mythos regularly since then, though not for cybersecurity purposes, said the person, who corroborated the account with screenshots and a live demonstration of the model.
The users relied on a mix of tactics to get into Mythos. These included using access the person had as a worker at a third-party contractor for Anthropic and trying commonly used internet sleuthing tools often employed by cybersecurity researchers, the person said. The users are part of a private Discord channel that focuses on hunting for information about unreleased models, including by using bots to scour for details that Anthropic and others have posted on unsecured websites such as GitHub.
“We’re investigating a report claiming unauthorized access to Claude Mythos Preview through one of our third-party vendor environments,” a spokesperson for Anthropic said in a statement.
Bloomberg.com
Anthropic’s Mythos Model Is Being Accessed by Unauthorized Users
A small group of unauthorized users have accessed Anthropic PBC’s new Mythos AI model, a technology that the company says is so powerful it can enable dangerous cyberattacks, according to a person familiar with the matter and documentation viewed by Bloomberg…
👍3🥰3💯2
MIT & the IMO released MathNet, the world’s largest dataset of International Math Olympiad problems & solutions.
MathNet is 5x larger than previous datasets & is sourced from over 40 countries across 4 decades.
MathNet is 5x larger than previous datasets & is sourced from over 40 countries across 4 decades.
Lobster Capital became the first agent-ready VC.
Team published an llms.txt, a structured file that AI agents (ChatGPT, Claude, Perplexity) read to understand who they are, what they invest in, and how to reach their.
Why it matters: founders and LPs increasingly research funds through AI. The VCs who don't show up in those answers won't get the meeting.
Team published an llms.txt, a structured file that AI agents (ChatGPT, Claude, Perplexity) read to understand who they are, what they invest in, and how to reach their.
Why it matters: founders and LPs increasingly research funds through AI. The VCs who don't show up in those answers won't get the meeting.
❤6🔥5💯1
Google launched Gemini Enterprise Agent Platform a platform for businesses to develop, scale, govern and optimize agents.
It’s the evolution of Vertex AI, bringing together model selection and agent building with new features for integration, security and more.
It gives access to 200+ of the world’s leading models through the Model Garden.
This includes Gemini 3.1 Pro, Gemini 3.1 Flash Image, and Lyria 3, alongside open models like Gemma 4.
It’s the evolution of Vertex AI, bringing together model selection and agent building with new features for integration, security and more.
It gives access to 200+ of the world’s leading models through the Model Garden.
This includes Gemini 3.1 Pro, Gemini 3.1 Flash Image, and Lyria 3, alongside open models like Gemma 4.
This media is not supported in your browser
VIEW IN TELEGRAM
Incredible work by Sony.They’ve built “Ace”, an autonomous ping-pong robot that uses RL and Sony’s vision sensors to achieve expert-level play in ping pong. A huge leap forward for adaptive robotics.
Ace robot beats 3 of 5 elite table tennis players. Loses to professionals.
Human players win points with faster-than-average shots (p<0.001 between won vs returned). Ace wins with ordinary shots. Same speed and spin profile whether it wins or loses the rally (p=0.88).
It's playing a completely different sport than the humans are.
Trained entirely in simulation. Zero sim-to-real tricks beyond good physics modeling and asymmetric actor-critic (critic sees ground truth, actor sees noisy sensors).
GitHub.
Ace robot beats 3 of 5 elite table tennis players. Loses to professionals.
Human players win points with faster-than-average shots (p<0.001 between won vs returned). Ace wins with ordinary shots. Same speed and spin profile whether it wins or loses the rally (p=0.88).
It's playing a completely different sport than the humans are.
Trained entirely in simulation. Zero sim-to-real tricks beyond good physics modeling and asymmetric actor-critic (critic sees ground truth, actor sees noisy sensors).
GitHub.
❤4🔥4👏2
Autonomous AI agents just found 10 zero-days in Chrome and the model almost didn’t matter
AgentFlow, a multi-agent harness developed by researchers at UCSB and Fuzzland, autonomously discovered 10 previously unknown vulnerabilities in Google Chrome over 7 days — including 2 Critical sandbox-escape CVEs confirmed by Google’s Vulnerability Reward Program.
The model powering the campaign was Kimi K2.5 managed to find and exploit 10 vulnerabilities in browsers.
AgentFlow used Kimi K2.5 not because it’s the best model, but because running 192 parallel agents for 7 days at Claude Opus prices was impractical and found 10 zero-days in Chrome, including 2 Critical sandbox-escape CVEs confirmed by Google VRP.
Sandbox escape is a serious vulnerability class that breaks Chrome’s process isolation, but turning it into a full system compromise requires a multi-step exploit chain.
The paper withholds the PoCs entirely.
The real story here: the harness architecture did most of the work.
The same framework, with Claude Opus 4.6, scored #1 on TerminalBench-2. The model matters less than how you orchestrate it.
AgentFlow, a multi-agent harness developed by researchers at UCSB and Fuzzland, autonomously discovered 10 previously unknown vulnerabilities in Google Chrome over 7 days — including 2 Critical sandbox-escape CVEs confirmed by Google’s Vulnerability Reward Program.
The model powering the campaign was Kimi K2.5 managed to find and exploit 10 vulnerabilities in browsers.
AgentFlow used Kimi K2.5 not because it’s the best model, but because running 192 parallel agents for 7 days at Claude Opus prices was impractical and found 10 zero-days in Chrome, including 2 Critical sandbox-escape CVEs confirmed by Google VRP.
Sandbox escape is a serious vulnerability class that breaks Chrome’s process isolation, but turning it into a full system compromise requires a multi-step exploit chain.
The paper withholds the PoCs entirely.
The real story here: the harness architecture did most of the work.
The same framework, with Claude Opus 4.6, scored #1 on TerminalBench-2. The model matters less than how you orchestrate it.
arXiv.org
Synthesizing Multi-Agent Harnesses for Vulnerability Discovery
LLM agents have begun to find real security vulnerabilities that human auditors and automated fuzzers missed for decades, in source-available targets where the analyst can build and instrument the...
❤4🔥3💯2
Google just proved Image generators are generalist vision learners
They introduced Vision Banana, a model built by instruction-tuning a base image generator (Nano Banana Pro).
Instead of using special systems for different tasks, they reframe every vision problem like segmentation or depth estimation as simply generating a new image. Think of it as drawing the answer instead of calculating it.
Vision Banana beats domain-specific experts, including the Segment Anything Model 3 (SAM 3) on segmentation and the Depth Anything series on metric depth estimation, all without sacrificing its original ability to create images.
This suggests generative pretraining is the new foundation for all of computer vision.
They introduced Vision Banana, a model built by instruction-tuning a base image generator (Nano Banana Pro).
Instead of using special systems for different tasks, they reframe every vision problem like segmentation or depth estimation as simply generating a new image. Think of it as drawing the answer instead of calculating it.
Vision Banana beats domain-specific experts, including the Segment Anything Model 3 (SAM 3) on segmentation and the Depth Anything series on metric depth estimation, all without sacrificing its original ability to create images.
This suggests generative pretraining is the new foundation for all of computer vision.
vision-banana.github.io
Vision Banana | Google DeepMind
A generalist model achieving state-of-the-art on segmentation, depth, and surface normal tasks.
🆒4❤3👏1💯1
Anthropic just introduced forked subagents in their latest update
Unlike regular subagents, forked subagents can inherit the same context as the main agent. This looks convenient for cases where richer context matters more.
Unlike regular subagents, forked subagents can inherit the same context as the main agent. This looks convenient for cases where richer context matters more.
Claude Code Docs
Claude Code changelog - Claude Code Docs
Release notes for Claude Code, including new features, improvements, and bug fixes by version.
🥰3🔥2👏2
DeepSeek-V4 Preview is officially live & open-sourced
DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Open weights.
DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Open weights.
👏3🔥2💯2
OMG 😯AI companies Cohere of Canada and Aleph Alpha of Germany have agreed to merge