A new Google DeepMind paper reveals a vulnerability in the AI supply chain
In this paper, "Cascading Adversarial Bias," shows how tiny, malicious changes to a large "teacher" language model can create amplified biases in smaller "student" models after distillation.
Many of today's efficient AI models are created through "distillation," where a smaller student model learns from a larger teacher. Researchers investigated what happens when an adversary poisons a bit of teacher's training data.
Model distillation, while vital for creating efficient AI, opens a dangerous pathway for undetectable and amplified bias. This represents a trustworthiness problem for the entire AI ecosystem.
In this paper, "Cascading Adversarial Bias," shows how tiny, malicious changes to a large "teacher" language model can create amplified biases in smaller "student" models after distillation.
Many of today's efficient AI models are created through "distillation," where a smaller student model learns from a larger teacher. Researchers investigated what happens when an adversary poisons a bit of teacher's training data.
Model distillation, while vital for creating efficient AI, opens a dangerous pathway for undetectable and amplified bias. This represents a trustworthiness problem for the entire AI ecosystem.
OpenAI is testing an experimental AI phone agent at 1-888-GPT-0090 that's designed to answer general questions about ChatGPT and other OpenAI products for callers from US or Canadian numbers.
OpenAI Help Center
1-888-GPT-0090 - AI voice help over the phone (experimental) | OpenAI Help Center
🦄3
Mistral is releasing its own “vibe coding” client, Mistral Code, to compete with incumbents like Windsurf, Anysphere’s Cursor, and GitHub Copilot.
Mistral Code, a fork of the open source project Continue, is an AI-powered coding assistant that bundles Mistral’s models, an “in-IDE” assistant, local deployment options, and enterprise tooling into a single package.
A private beta is available as of Wednesday for JetBrains development platforms and Microsoft’s VS Code.
At its core, Mistral Code is powered by four models that are state of the art in coding:
- Codestral for fill-in-the-middle / code autocomplete
- Codestral Embed for code search and retrieval
- Devstral for agentic coding
- And Mistral Medium for chat assistance
Mistral Code, a fork of the open source project Continue, is an AI-powered coding assistant that bundles Mistral’s models, an “in-IDE” assistant, local deployment options, and enterprise tooling into a single package.
A private beta is available as of Wednesday for JetBrains development platforms and Microsoft’s VS Code.
At its core, Mistral Code is powered by four models that are state of the art in coding:
- Codestral for fill-in-the-middle / code autocomplete
- Codestral Embed for code search and retrieval
- Devstral for agentic coding
- And Mistral Medium for chat assistance
Mistral AI
Introducing Mistral Code | Mistral AI
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
🔥8
OpenAI livestream "Announces updates to ChatGPT for business"
OpenAI
Livestreams
Watch OpenAI livestreams featuring product launches, demos, and announcements. Explore past events, discover the latest AI innovations, and stay up to date.
All about AI, Web 3.0, BCI
OpenAI livestream "Announces updates to ChatGPT for business"
OpenAI is dropping custom and ready-to-use data connectors for ChatGPT, including Google Drive, Box, Hubspot, and Outlook
The co also added a "record mode" to help teams turn meetings into transcriptions with key points, action items, or plans.
The co also added a "record mode" to help teams turn meetings into transcriptions with key points, action items, or plans.
Agentic_AI_for_Intelligent_Business_Operations__1749123645.pdf
1.5 MB
Agentic AI for Intelligent Business Operations. IBM has released an insightful research report on the potential transformational effect on organizations that are turning to agentic AI to capture a competitive edge.
Key Takeaways of the Research:
1. Agentic AI enables business operations to autonomously learn, adapt, and optimize in real time, offering more than speed—it drives proactive #innovation and personalization.
2. The shift is from automating isolated tasks to orchestrating dynamic, self-improving processes through intelligent agents.
3. Successful AI deployment hinges on integrating people and technology, where human oversight, decision-making, and creativity remain essential.
4. IBM IBV research shows over 80% of executives view global business service automation as a strategic priority, expecting AI agents to lead the transformation.
5. By 2027, 86% of executives believe AI agents will significantly enhance process automation and workflow redesign.
6. Agentic AI is reshaping the workplace—tech manages operations while talent manages tech, with many users already interacting primarily through AI assistants.
7. Currently, 76% of organizations are piloting or scaling autonomous AI agents to manage intelligent workflows.
8. AI agents operate on goals and context rather than rigid rules, dynamically adjusting to achieve outcomes—similar to how autonomous vehicles function.
9. 75% of executives expect AI agents to handle transactional workflows autonomously within the next two years.
Key Takeaways of the Research:
1. Agentic AI enables business operations to autonomously learn, adapt, and optimize in real time, offering more than speed—it drives proactive #innovation and personalization.
2. The shift is from automating isolated tasks to orchestrating dynamic, self-improving processes through intelligent agents.
3. Successful AI deployment hinges on integrating people and technology, where human oversight, decision-making, and creativity remain essential.
4. IBM IBV research shows over 80% of executives view global business service automation as a strategic priority, expecting AI agents to lead the transformation.
5. By 2027, 86% of executives believe AI agents will significantly enhance process automation and workflow redesign.
6. Agentic AI is reshaping the workplace—tech manages operations while talent manages tech, with many users already interacting primarily through AI assistants.
7. Currently, 76% of organizations are piloting or scaling autonomous AI agents to manage intelligent workflows.
8. AI agents operate on goals and context rather than rigid rules, dynamically adjusting to achieve outcomes—similar to how autonomous vehicles function.
9. 75% of executives expect AI agents to handle transactional workflows autonomously within the next two years.
❤5
FutureHouse released ether0, the first scientific reasoning model
Researchers trained Mistral 24B with RL on several molecular design tasks in chemistry.
Remarkably, researchers found that LLMs can learn some scientific tasks more much data-efficiently than specialized models trained from scratch on the same data, and can greatly outperform frontier models and humans on those tasks.
For at least a subset of scientific classification, regression, and generation problems, post-training LLMs may provide a much more data-efficient approach than traditional machine learning approaches.
Paper.
Weights
Researchers trained Mistral 24B with RL on several molecular design tasks in chemistry.
Remarkably, researchers found that LLMs can learn some scientific tasks more much data-efficiently than specialized models trained from scratch on the same data, and can greatly outperform frontier models and humans on those tasks.
For at least a subset of scientific classification, regression, and generation problems, post-training LLMs may provide a much more data-efficient approach than traditional machine learning approaches.
Paper.
Weights
www.futurehouse.org
ether0: a scientific reasoning model for chemistry | FutureHouse
Epoch AI just published a dataset of worldwide AI supercomputers (GPU clusters)!
Epoch AI
Data on GPU clusters
Our database of over 500 GPU clusters and supercomputers tracks large hardware facilities, including those used for AI training and inference.
Qwen just dropped its new embedding model
Qwen3-Embedding offers a range of sizes (0.6B, 4B, 8B) for text embedding & reranking, achieving SOTA performance on MTEB.
Qwen3-Embedding offers a range of sizes (0.6B, 4B, 8B) for text embedding & reranking, achieving SOTA performance on MTEB.
Qwen
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
GITHUB HUGGING FACE MODELSCOPE DISCORD
We release Qwen3 Embedding series, a new proprietary model of the Qwen model family. These models are specifically designed for text embedding, retrieval, and reranking tasks, built on the Qwen3 foundation model. Leveraging…
We release Qwen3 Embedding series, a new proprietary model of the Qwen model family. These models are specifically designed for text embedding, retrieval, and reranking tasks, built on the Qwen3 foundation model. Leveraging…
👏3💅2
Boltz-2 is a new biomolecular foundation model that goes beyond AlphaFold3 and Boltz-1 by jointly modeling complex structures and binding affinities, a critical component towards accurate molecular design.
Boltz-2 a new model capable not only of predicting structures but also binding affinities
Boltz-2 is the first AI model to approach the performance of FEP simulations while being more than 1000x faster.
All open-sourced under MIT license.
Paper
Boltz-2 a new model capable not only of predicting structures but also binding affinities
Boltz-2 is the first AI model to approach the performance of FEP simulations while being more than 1000x faster.
All open-sourced under MIT license.
Paper
GitHub
GitHub - jwohlwend/boltz: Official repository for the Boltz biomolecular interaction models
Official repository for the Boltz biomolecular interaction models - jwohlwend/boltz
👍6
Extract an AI system developed by the UK Government’s AI Incubator team using Google’s Gemini model
It aims to modernize the UK’s planning system by converting old, paper-based planning documents—such as blurry maps and handwritten notes—into clear, digital data in about 40 seconds.
Key points:
1. Speeds up processing of ~350,000 annual planning applications in England, supporting housing and infrastructure development.
2. How it works:
- Uses Gemini’s multimodal reasoning to extract critical information from text, handwritten notes, and low-quality map images.
- Identifies map features (e.g., boundaries, shaded areas) using tools like OpenCV, Ordnance Survey, and Segment Anything.
- Matches historical maps to modern equivalents using addresses, landmarks, and feature mapping (e.g., LoFTR) to convert shapes into precise geographical coordinates.
- Reduces council workload, simplifies processes, and frees staff for strategic planning. It supports the UK’s goal of building 1.5 million new homes.
It aims to modernize the UK’s planning system by converting old, paper-based planning documents—such as blurry maps and handwritten notes—into clear, digital data in about 40 seconds.
Key points:
1. Speeds up processing of ~350,000 annual planning applications in England, supporting housing and infrastructure development.
2. How it works:
- Uses Gemini’s multimodal reasoning to extract critical information from text, handwritten notes, and low-quality map images.
- Identifies map features (e.g., boundaries, shaded areas) using tools like OpenCV, Ordnance Survey, and Segment Anything.
- Matches historical maps to modern equivalents using addresses, landmarks, and feature mapping (e.g., LoFTR) to convert shapes into precise geographical coordinates.
- Reduces council workload, simplifies processes, and frees staff for strategic planning. It supports the UK’s goal of building 1.5 million new homes.
Google
UK government harnesses Gemini to support faster planning decisions
Extract, built with Gemini, uses the model’s advanced visual reasoning and multi-modal capabilities to help councils turn old planning documents—including blurry maps and handwritten notes—into clear, digital data, speeding up decision-making timelines for…
👏3
Project_Pine_Tokenised_Financial_Markets_1749471407.pdf
1.7 MB
Project Pine - Tokenised Financial Markets by The Federal Reserve Bank of New York and BIS
ProjectPine found that central banks could customise and deploy policy implementation tools using programmable smart contracts in a potential future state where commercial banks and other private sector financial institutions have widely adopted tokenisation for wholesale payments and securities settlement.
The project generated the prototype of a generic monetary policy implementation tokenised toolkit for potential further research and development by central banks across jurisdictions and currencies. The prototype was designed to be technically modifiable for different central banks' monetary policy frameworks and calibrated to conduct standard or emergency market operations.
The toolkit prototype was created in consultation with central banks' financial markets advisors from multiple jurisdictions, who helped outline the project scope and specific design requirements. It is not particular to any currency or jurisdiction. It can fulfil a common set of central bank implementation requirements, including paying interest on reserves, open market operations, and collateral management.
ProjectPine found that central banks could customise and deploy policy implementation tools using programmable smart contracts in a potential future state where commercial banks and other private sector financial institutions have widely adopted tokenisation for wholesale payments and securities settlement.
The project generated the prototype of a generic monetary policy implementation tokenised toolkit for potential further research and development by central banks across jurisdictions and currencies. The prototype was designed to be technically modifiable for different central banks' monetary policy frameworks and calibrated to conduct standard or emergency market operations.
The toolkit prototype was created in consultation with central banks' financial markets advisors from multiple jurisdictions, who helped outline the project scope and specific design requirements. It is not particular to any currency or jurisdiction. It can fulfil a common set of central bank implementation requirements, including paying interest on reserves, open market operations, and collateral management.
🔥3
Alibaba's RL LLM training library: ROLL
ROLL is built upon several key modules to serve these user groups effectively:
1. A single-controller architecture combined with an abstraction of the parallel worker simplifies the development of the training pipeline.
2. The parallel strategy and data transfer modules enable efficient and scalable training.
3. The rollout scheduler offers fine-grained management of each sample's lifecycle during the rollout stage.
4. The environment worker and reward worker support rapid and flexible experimentation with agentic RL algorithms and reward designs.
Finally, AutoDeviceMapping allows users to assign resources to different models flexibly across various stages.
GitHub.
ROLL is built upon several key modules to serve these user groups effectively:
1. A single-controller architecture combined with an abstraction of the parallel worker simplifies the development of the training pipeline.
2. The parallel strategy and data transfer modules enable efficient and scalable training.
3. The rollout scheduler offers fine-grained management of each sample's lifecycle during the rollout stage.
4. The environment worker and reward worker support rapid and flexible experimentation with agentic RL algorithms and reward designs.
Finally, AutoDeviceMapping allows users to assign resources to different models flexibly across various stages.
GitHub.
Autonomous Agents That Think, Remember, and Evolve from AWS team
This project showcasing how autonomous agents can move beyond simple tasks to reason, remember, and adapt
- powered by Mem0, AWS,, and Strands Agents SDK
From remembering past findings to adapting strategies on the fly, Cyber-AutoAgent uses long- and short-term memory to build real expertise - one interaction at a time.
GitHub.
This project showcasing how autonomous agents can move beyond simple tasks to reason, remember, and adapt
- powered by Mem0, AWS,, and Strands Agents SDK
From remembering past findings to adapting strategies on the fly, Cyber-AutoAgent uses long- and short-term memory to build real expertise - one interaction at a time.
GitHub.
At WWDC Apple introduced a new generation of LLMs developed to enhance the Apple Intelligence features.
Also introduced the new Foundation Models framework, which gives app developers direct access to the on-device foundation language model.
Also introduced the new Foundation Models framework, which gives app developers direct access to the on-device foundation language model.
Apple Machine Learning Research
Updates to Apple’s On-Device and Server Foundation Language Models
With Apple Intelligence, we're integrating powerful generative AI right into the apps and experiences people use every day, all while…
GALBOT announced OpenWBT – an open-source, whole-body humanoid VR teleoperation system using Apple Vision Pro.
It supports Unitree G1 and H1 robots, enabling operators to control movements like walking, squatting, bending, grasping, and lifting.
It supports Unitree G1 and H1 robots, enabling operators to control movements like walking, squatting, bending, grasping, and lifting.
GitHub
GitHub - GalaxyGeneralRobotics/OpenWBT: Official implementation of OpenWBT.
Official implementation of OpenWBT. Contribute to GalaxyGeneralRobotics/OpenWBT development by creating an account on GitHub.
👏3
Microsoft introduced Code Researcher - a deep research agent for large systems code and commit history.
Achieves a 58% crash resolution rate on a benchmark of crashes in the Linux kernel, a complex codebase with 28M LOC & 75K files.
Achieves a 58% crash resolution rate on a benchmark of crashes in the Linux kernel, a complex codebase with 28M LOC & 75K files.
Microsoft Research
Code Researcher: Deep Research Agent for Large Systems Code and Commit History - Microsoft Research
Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code remains underexplored. Due to the size and complexities of systems code, making changes to a systems codebase is a daunting…
👍3🦄3