AI & ML Papers
Photo
🔥 DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
📅 Published on Mar 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2603.26164
• PDF: https://arxiv.org/pdf/2603.26164
• Project Page: https://opendcai.github.io/DataFlex-Doc/en/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#LargeLanguageModels #DataCentricTraining #DynamicTrainingMethods #LanguageModelOptimization #DataDrivenAI
💡 The paper introduces DataFlex, a unified framework for dynamic data-centric training of large language models. The problem addressed is that existing approaches to data selection, data mixture optimization, and data reweighting are often developed in isolated codebases, making it difficult to reproduce, compare, and integrate them. DataFlex solves this problem by providing a unified framework that supports three major paradigms of dynamic data optimization: sample selection, domain mixture adjustment, and sample reweighting.
The method involves building DataFlex upon the LLaMA-Factory framework, which allows for extensible trainer abstractions and modular components. This enables a drop-in replacement for standard large language model training and unifies key model-dependent operations such as embedding extraction, inference, and gradient computation. DataFlex is also compatible with large-scale settings, including DeepSpeed ZeRO-3.
The results show that DataFlex provides an effective, efficient, and reproducible infrastructure for data-centric dynamic training of large language models. Comprehensive experiments demonstrate that dynamic data selection consistently outperforms static full-data training, and data mixture methods improve both accuracy and perplexity over default proportions. Additionally, DataFlex achieves consistent runtime improvements over original implementations. Overall, the paper contributes a unified framework that enables efficient large-scale deployment of data-centric dynamic training methods for large language models.
📅 Published on Mar 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2603.26164
• PDF: https://arxiv.org/pdf/2603.26164
• Project Page: https://opendcai.github.io/DataFlex-Doc/en/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#LargeLanguageModels #DataCentricTraining #DynamicTrainingMethods #LanguageModelOptimization #DataDrivenAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
📅 Published on Jul 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.09635
• PDF: https://arxiv.org/pdf/2605.09635
• Project Page: https://haolpku.github.io/K12-KGraph-page/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/lhpku20010120/K12-KGraph
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#EducationalLanguageModels #CurriculumAlignedKnowledgeGraph #LargeLanguageModels #KnowledgeGraphEmbeddings #BenchmarkingEducationalAI
💡 The paper introduces K12-KGraph, a curriculum-aligned knowledge graph for benchmarking and training educational large language models. The existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented, which is referred to as curriculum cognition. K12-KGraph is extracted from official textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school, containing nine node types and fourteen relation types covering curriculum structure and visual grounding.
From this graph, the authors derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. They also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs.
The results show that existing large language models, such as Gemini-3-Flash and Gemma-4-31B-IT, achieve only 57 percent and 46 percent exact match on K12-Bench, with Prereq and Neighbor being the hardest tasks. The training experiments demonstrate that domain-specific supervision can reduce this gap. Under an unmatched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaoKao Bench and EdEval. For vision-language models, K12-Train-Full achieves the best overall results on GaoKao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full Data Flow and Wizard LM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary.
The authors release the graph, benchmark, training data, and complete construction pipeline, providing a valuable resource for developing and evaluating educational large language models. The paper's contributions include the creation of a curriculum-aligned knowledge graph, a comprehensive benchmark, and a graph-guided supervised fine-tuning corpus, which can help improve the performance of large language models in educational settings.
📅 Published on Jul 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2605.09635
• PDF: https://arxiv.org/pdf/2605.09635
• Project Page: https://haolpku.github.io/K12-KGraph-page/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/lhpku20010120/K12-KGraph
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#EducationalLanguageModels #CurriculumAlignedKnowledgeGraph #LargeLanguageModels #KnowledgeGraphEmbeddings #BenchmarkingEducationalAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤1
AI & ML Papers
Photo
🔥 Kimi K3: Open Frontier Intelligence
📅 Published on Jul 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.24653
• PDF: https://arxiv.org/pdf/2607.24653
• Project Page: https://www.kimi.com/blog/kimi-k3
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MixtureOfExpertsModel #VisionCapabilitiesInAI #LargeLanguageModels #AttentionMechanismsInDeepLearning #ScalingEfficiencyInAI
💡 The paper introduces Kimi K3, a 2.8 trillion parameter mixture of experts model with native vision capabilities and a 1 million token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. The model also incorporates Stable Latent MoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes. These advances yield an approximately 2.5 times improvement in overall scaling efficiency over Kimi K2.
The model achieves frontier level performance across long horizon coding, genetic, knowledge, reasoning, and vision tasks. Although its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT 5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in the study.
The key contributions of the paper include the introduction of Kimi K3, which is supported by infrastructure advances in multiple areas, such as algorithm system co design for KD, perfectly balanced expert parallel training with efficient memory management, million token genetic RL with persistent rollout and sandbox states, and deployment innovations. The full Kimi K3 model weights are released to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.
The problem addressed in the paper is the development of a highly efficient and scalable model that can achieve state of the art performance across a wide range of tasks. The method used to address this problem is the introduction of Kimi K3, which incorporates several key innovations, including Kimi Delta Attention, Attention Residuals, and Stable Latent MoE. The results of the study demonstrate the effectiveness of Kimi K3, which achieves frontier level performance across a range of tasks and outperforms other open and proprietary models. Overall, the paper contributes to the development of highly efficient and scalable models that can achieve state of the art performance across a wide range of tasks.
📅 Published on Jul 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.24653
• PDF: https://arxiv.org/pdf/2607.24653
• Project Page: https://www.kimi.com/blog/kimi-k3
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MixtureOfExpertsModel #VisionCapabilitiesInAI #LargeLanguageModels #AttentionMechanismsInDeepLearning #ScalingEfficiencyInAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
🔥 JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23588
• PDF: https://arxiv.org/pdf/2607.23588
• Project Page: https://www.jarvishub.site/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MultimodalCreativeAgents #CanvasNativeGeneration #GenerativeModeling #CreativeArtificialIntelligence #MultimodalProductionSystems
💡 The paper introduces JarvisHub, an open harness for canvas-native multimodal creative agents, which aims to address the limitations of existing generative models in supporting long-horizon multimodal creative production. Current models can synthesize high-quality images, videos, audio clips, and other creative assets, but real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state.
Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift towards agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time.
JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use towards sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.
The paper's contributions include the introduction of JarvisHub as an open harness for canvas-native multimodal creative agents, which provides a flexible and inspectable framework for long-horizon multimodal creative production. The system allows agents to act within an editable canvas, representing the user workspace, agent's external memory, action space, and shared project state. The three-layer architecture of JarvisHub enables agents to progressively plan, generate, revise, and organize multimodal projects, while users can inspect, guide, and intervene throughout the process. Overall, JarvisHub has the potential to advance the field of creative AI by providing a more flexible and sustainable approach to multimodal creative production.
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23588
• PDF: https://arxiv.org/pdf/2607.23588
• Project Page: https://www.jarvishub.site/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MultimodalCreativeAgents #CanvasNativeGeneration #GenerativeModeling #CreativeArtificialIntelligence #MultimodalProductionSystems
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 Data Pyramid for Embodied Manipulation
📅 Published on Jul 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.24744
• PDF: https://arxiv.org/pdf/2607.24744
• Project Page: https://jasper-aaa.github.io/embodied-data-pyramid/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#EmbodiedManipulation #RobotLearning #MultimodalData #EmbodiedAI #FoundationModels
💡 The paper proposes a data pyramid framework for embodied manipulation, which organizes data sources into five complementary categories: real-robot data, UMl-style data, egocentric and exocentric data, simulation data, and general vision-language data. The framework is designed to address the challenges of embodied agents, which require data that couples observations with physical states and actions. The authors analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during training. They relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. The paper also discusses six open challenges, including building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. The proposed data pyramid framework provides a foundation for the design of next-generation embodied systems, which can learn to perceive, reason, and act in complex environments. The authors hope that this work will pave the way for the development of more advanced embodied agents that can learn from diverse data sources and improve their performance in various tasks.
📅 Published on Jul 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.24744
• PDF: https://arxiv.org/pdf/2607.24744
• Project Page: https://jasper-aaa.github.io/embodied-data-pyramid/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#EmbodiedManipulation #RobotLearning #MultimodalData #EmbodiedAI #FoundationModels
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23855
• PDF: https://arxiv.org/pdf/2607.23855
• Project Page: https://openmoss.ai/OmniVAE.github.io/
🤖 Models citing this paper:
• https://huggingface.co/OpenMOSS-Team/OmniVAE
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#AudioVideoGeneration #CrossModalAlignment #VariationalAutoencoders #MultimodalLearning #JointGenerationModels
💡 The paper introduces OmniVAE, a joint audio-video variational autoencoder that learns fine-grained semantic alignment between audio and video latent representations. Recent generative models have moved beyond silent video or standalone audio synthesis towards joint generation of synchronized audio and video, but this remains challenging due to the fundamental structural differences between the two modalities. Most existing methods use audio and video VAEs trained separately, resulting in a lack of cross-modal alignment, which leaves the downstream generative model to learn cross-modal synchronization from scratch.
OmniVAE addresses this issue by jointly training an audio-video VAE that captures temporal-semantic correspondence and aligns the two latent spaces. It uses a segment-level audio-video contrastive objective to achieve this alignment and also distills features from pre-trained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces.
The results of extensive experiments show that OmniVAE consistently improves the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation tasks. The findings underscore the importance of learning unified representations as a foundation for omnimodal modeling. Overall, OmniVAE provides a novel approach to joint audio-video generation, enabling more effective and synchronized generation of audio and video content.
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23855
• PDF: https://arxiv.org/pdf/2607.23855
• Project Page: https://openmoss.ai/OmniVAE.github.io/
🤖 Models citing this paper:
• https://huggingface.co/OpenMOSS-Team/OmniVAE
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#AudioVideoGeneration #CrossModalAlignment #VariationalAutoencoders #MultimodalLearning #JointGenerationModels
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 GNM Head: A Generative aNthropometric Model of the human head
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23687
• PDF: https://arxiv.org/pdf/2607.23687
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#GenerativeModeling #AnthropometricAnalysis #3DHeadModeling #FacialReconstruction #HumanComputerInteraction
💡 The paper introduces a new parametric model called the Generative Anthropometric Model of the human head, or GNM Head. The model is designed to address the limitations of existing publicly available models, which typically only capture the outer geometry of the head and ignore internal structures such as the eyes and mouth. These existing models also often suffer from reduced geometric quality due to low-fidelity input data sets.
The GNM Head model is built on an extensive database of high-resolution 3D scans combined with high-quality anatomy-specific artist-made samples. The model encompasses the head, face, neck, eyeballs, teeth, and tongue, and includes specialized sub-models for the ocular and intra-oral structures.
The paper details the data provenance, model architecture, and performance of the GNM Head model, including its ability to fit target 3D face scans. The results show that the GNM Head model is a significant improvement over existing models, offering a more comprehensive and accurate representation of the human head.
To foster community innovation, the complete GNM framework is made publicly available. The GNM Head model has the potential to serve as a crucial conditioning signal within generative large vision models, allowing for tight spatial control of generated imagery. Overall, the paper contributes a new and improved parametric model of the human head, which can be used in a variety of applications, including computer vision, graphics, and animation.
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23687
• PDF: https://arxiv.org/pdf/2607.23687
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#GenerativeModeling #AnthropometricAnalysis #3DHeadModeling #FacialReconstruction #HumanComputerInteraction
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤2
🔥 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13285
• PDF: https://arxiv.org/pdf/2607.13285
• Project Page: https://ruhan-wang.github.io/Harness-Handbook/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#EvolvingAgentHarnesses #ArtificialIntelligenceSystems #HarnessHandbook #BehaviorCentricDesign #AIModelMaintenance
💡 The paper discusses the challenges of evolving agent harnesses, which are critical components of modern AI systems. As models, APIs, environments, and requirements change, harnesses must be continuously modified. However, this is difficult because production harnesses are large, tightly coupled, and behaviorally distributed, making it hard to identify the code locations that implement specific behaviors.
The authors introduce the Harness Handbook, a behavior-centric representation that is synthesized automatically from a harness code base via static analysis and LLM-assisted structuring. The Handbook links each behavior to its corresponding source code, making it easier to navigate and edit the harness.
The authors also introduce Behavior-Guided Progressive Disclosure BGPD, a method that guides agents from high-level behaviors to relevant implementation details and verifies candidate locations against the current source code.
The results show that Handbook-Assisted planning improves behavior localization and edit-plan quality while using fewer planner tokens. The largest gains are seen on scattered sites, rarely executed paths, and cross-module interactions. The paper concludes that evolving complex agent systems depends not only on generating edits but also on determining where those edits should be made.
Overall, the paper provides a solution to the problem of evolving agent harnesses by introducing a behavior-centric representation and a method for guiding agents to relevant implementation details. The results demonstrate the effectiveness of the proposed approach in improving behavior localization and edit-plan quality.
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13285
• PDF: https://arxiv.org/pdf/2607.13285
• Project Page: https://ruhan-wang.github.io/Harness-Handbook/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#EvolvingAgentHarnesses #ArtificialIntelligenceSystems #HarnessHandbook #BehaviorCentricDesign #AIModelMaintenance
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤2
AI & ML Papers
Photo
🔥 Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
📅 Published on Jul 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.24904
• PDF: https://arxiv.org/pdf/2607.24904
• Project Page: https://microsoft.github.io/Mage
🤖 Models citing this paper:
• https://huggingface.co/microsoft/Mage-VL
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/microsoft/mage-vl-demo
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MultimodalFoundationModels #CodecNativeStreaming #VisionLanguageModels #RealTimeMultimodalUnderstanding #EfficientMultimodalProcessing
💡 The paper presents Mage-VL, an efficient codec-native streaming multimodal foundation model for real-time multimodal understanding and interaction. Standard vision-language models suffer from Moravec's paradox, exceling at complex offline visual reasoning but struggling with simple streaming perception tasks and processing them inefficiently. To address this, the authors propose a custom tokenizer, Mage-ViT, which replaces uniform frame sampling with selective encoding of dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor and predicted frames. This approach reduces visual token consumption by over 75 percent while preserving spatiotemporal context.
Mage-ViT is trained from scratch on approximately 560 million unlabeled images and 100 million unlabeled video frames, and it matches or outperforms flagship encoders trained on billions of image-text pairs. The authors also establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.
Furthermore, through a bio-inspired dual-system architecture, consisting of a lightweight System 1 event gate and a causal System 2 decoder, Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5 times wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline.
The paper delivers seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, Video QA SF redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL. Overall, the paper presents a novel approach to multimodal understanding and interaction, with significant improvements in efficiency, accuracy, and speed.
📅 Published on Jul 27
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.24904
• PDF: https://arxiv.org/pdf/2607.24904
• Project Page: https://microsoft.github.io/Mage
🤖 Models citing this paper:
• https://huggingface.co/microsoft/Mage-VL
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/microsoft/mage-vl-demo
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MultimodalFoundationModels #CodecNativeStreaming #VisionLanguageModels #RealTimeMultimodalUnderstanding #EfficientMultimodalProcessing
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 Pass the Baton: Trajectory-Relayed On-Policy Distillation
📅 Published on Jul 28
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.26057
• PDF: https://arxiv.org/pdf/2607.26057
• Project Page: https://zju-real.github.io/Relay-OPD/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#OnPolicyDistillation #TrajectoryRelayedLearning #PrefixFailureMitigation #TeacherStudentContinuationAsymmetry #RelayOnPolicyDistillation
💡 The paper addresses the issue of prefix failure in on-policy distillation, where the student model commits to a wrong reasoning direction and subsequent generations build on this deviation, producing misguided continuations. The authors identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction. They propose Relay On-Policy Distillation, which converts this asymmetry into a label-free handoff trigger.
During training, Relay On-Policy Distillation constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy.
The method is evaluated on eight mathematical reasoning benchmarks with a Qwen 3-4B-Instruction-2507 teacher and Qwen 3-0.6B/1.7B-Non-Thinking students. The results show that Relay On-Policy Distillation achieves the best or second-best results on every benchmark, outperforming standard on-policy distillation by 5.73 percent and the strongest baseline Fast On-Policy Distillation by 1.49 percent on average for 1.7B, with consistent gains at 0.6B. Additionally, the training trajectory length is reduced by over 50 percent.
Overall, the paper proposes a novel method to address prefix failure in on-policy distillation, achieving significant improvements in performance and efficiency.
📅 Published on Jul 28
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.26057
• PDF: https://arxiv.org/pdf/2607.26057
• Project Page: https://zju-real.github.io/Relay-OPD/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#OnPolicyDistillation #TrajectoryRelayedLearning #PrefixFailureMitigation #TeacherStudentContinuationAsymmetry #RelayOnPolicyDistillation
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
📅 Published on Jul 28
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.25948
• PDF: https://arxiv.org/pdf/2607.25948
• Project Page: https://modus-multimodal.epfl.ch/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MultimodalLearning #DecoderOnlyArchitectures #AnyToAnyModeling #ModalityTransfer #CrossModalPrediction
💡 The paper introduces MODUS, a novel any-to-any modeling approach that enables the prediction of any modality from any combination of others within a single network. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, which can impact their performance and prevent them from leveraging strong pre-trained decoder-only models. In contrast, MODUS treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. This is achieved through a decoder-only architecture that allows every modality to be both an input and an output of the same model. The resulting model, MODUS, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. The authors demonstrate that MODUS exhibits strong out-of-the-box performance and is competitive with specialist and multi-task baseline models using a single model across various benchmarks. Overall, MODUS provides a flexible and effective approach to any-to-any modeling, and all materials are open-sourced for further research and development.
📅 Published on Jul 28
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.25948
• PDF: https://arxiv.org/pdf/2607.25948
• Project Page: https://modus-multimodal.epfl.ch/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/epfl-vilab-modus/MODUS-15Modality
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://t.iss.one/PaperNexus
#MultimodalLearning #DecoderOnlyArchitectures #AnyToAnyModeling #ModalityTransfer #CrossModalPrediction
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤1