07. IA agentica
updated
Paper
• 2606.01533
• Published • 7
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper
• 2606.06741
• Published • 29
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper
• 2606.07412
• Published • 12
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper
• 2606.08348
• Published • 16
SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating
Paper
• 2606.07074
• Published • 12
Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
Paper
• 2606.06523
• Published • 7
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Paper
• 2606.10917
• Published • 76
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts
Paper
• 2606.05922
• Published • 70
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Paper
• 2606.11025
• Published • 41
Rethinking the Divergence Regularization in LLM RL
Paper
• 2606.09821
• Published • 34
Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval
Paper
• 2606.04391
• Published • 11
What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems
Paper
• 2606.05304
• Published • 5
Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking
Paper
• 2606.07689
• Published • 6
Paper
• 2606.10650
• Published • 10
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
Paper
• 2606.12191
• Published • 71
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Paper
• 2606.11926
• Published • 130
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
Paper
• 2606.13120
• Published • 4
WebChallenger: A Reliable and Efficient Generalist Web Agent
Paper
• 2606.10423
• Published • 2
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Paper
• 2606.12370
• Published • 21
Large Language Models Are Overconfident in Their Own Responses
Paper
• 2606.03437
• Published • 3
HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
Paper
• 2606.12882
• Published • 15
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
Paper
• 2606.10747
• Published • 13
LLM Agents Can See Code Repositories
Paper
• 2606.14061
• Published • 20
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
Paper
• 2606.06574
• Published • 25
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
Paper
• 2606.14249
• Published • 50
APPO: Agentic Procedural Policy Optimization
Paper
• 2606.12384
• Published • 78
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Paper
• 2606.15300
• Published • 13
FastContext: Training Efficient Repository Explorer for Coding Agents
Paper
• 2606.14066
• Published • 95
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Paper
• 2606.22388
• Published • 96
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Paper
• 2606.21337
• Published • 75
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
Paper
• 2606.22883
• Published • 37
Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
Paper
• 2606.18844
• Published • 20
Self-Compacting Language Model Agents
Paper
• 2606.23525
• Published • 19
Causal Discovery in the Era of Agents
Paper
• 2606.23608
• Published • 8
FastMix: Fast Data Mixture Optimization via Gradient Descent
Paper
• 2606.14971
• Published • 3
DanceOPD: On-Policy Generative Field Distillation
Paper
• 2606.27377
• Published • 81
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
Paper
• 2606.26790
• Published • 57
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
Paper
• 2606.26300
• Published • 53
PhysiFormer: Learning to Simulate Mechanics in World Space
Paper
• 2606.27364
• Published • 11
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
Paper
• 2606.27288
• Published • 4
Are We Ready For An Agent-Native Memory System?
Paper
• 2606.24775
• Published • 133
Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
Paper
• 2606.22953
• Published • 1
Forecasting Future Behavior as a Learning Task
Paper
• 2606.11445
• Published • 1
Qwen-AgentWorld: Language World Models for General Agents
Paper
• 2606.24597
• Published • 158
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
Paper
• 2606.24428
• Published • 52
OpenThoughts-Agent: Data Recipes for Agentic Models
Paper
• 2606.24855
• Published • 48
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Paper
• 2606.27378
• Published • 60
Towards Automating Scientific Review with Google's Paper Assistant Tool
Paper
• 2606.28277
• Published • 11
Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving
Paper
• 2606.27457
• Published • 4
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Paper
• 2607.01071
• Published • 31
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
Paper
• 2607.00461
• Published • 28
CausalMix: Data Mixture as Causal Inference for Language Model Training
Paper
• 2607.01104
• Published • 21
The State-Prediction Separation Hypothesis
Paper
• 2607.01218
• Published • 12
AutoTrainess: Teaching Language Models to Improve Language Models Autonomously
Paper
• 2606.31551
• Published • 24
Valdi: Value Diffusion World Models
Paper
• 2607.00917
• Published • 15
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Paper
• 2607.01211
• Published • 13
Autonomous Scientific Discovery via Iterative Meta-Reflection
Paper
• 2607.01131
• Published • 8
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
Paper
• 2606.32029
• Published • 14
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
Paper
• 2606.29082
• Published • 42
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
Paper
• 2606.23127
• Published • 26
Little Brains, Big Feats: Exploring Compact Language Models
Paper
• 2606.30062
• Published • 16
Hierarchical Experimentalist Agents
Paper
• 2606.29315
• Published • 5
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Paper
• 2607.01874
• Published • 23
PACE: A Proxy for Agentic Capability Evaluation
Paper
• 2607.02032
• Published • 20
When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search
Paper
• 2606.27669
• Published • 16
Paper
• 2607.27201
• Published • 104
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper
• 2608.05987
• Published • 94
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Paper
• 2608.05102
• Published • 65
Weak-to-Strong On-Policy Distillation
Paper
• 2607.26246
• Published • 56
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Paper
• 2608.05000
• Published • 59
Progressive Agent Skill Generation via Reinforcement Learning
Paper
• 2608.01678
• Published • 58
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
Paper
• 2608.06301
• Published • 34
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Paper
• 2608.02603
• Published • 34
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Paper
• 2608.05013
• Published • 35
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Paper
• 2608.03764
• Published • 27
Continual Learning in Transition
Paper
• 2608.06216
• Published • 23
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
Paper
• 2608.00155
• Published • 25
CADENA: Stepwise CAD Reverse Engineering
Paper
• 2608.00799
• Published • 40
Kimi K3: Open Frontier Intelligence
Paper
• 2607.24653
• Published • 491
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Paper
• 2607.02512
• Published • 309
Metis: Memory Foundation Model
Paper
• 2607.26760
• Published • 271
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Paper
• 2607.13285
• Published • 232
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Paper
• 2607.21461
• Published • 154
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Paper
• 2607.16617
• Published • 143
Weak-to-Strong Generalization via Direct On-Policy Distillation
Paper
• 2607.05394
• Published • 143
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
Paper
• 2607.17250
• Published • 93
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Paper
• 2607.05382
• Published • 88
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Paper
• 2607.04438
• Published • 64
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
Paper
• 2608.07222
• Published • 9
Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
Paper
• 2608.00675
• Published • 9
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
Paper
• 2608.06714
• Published • 9
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
Paper
• 2608.06501
• Published • 3
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
Paper
• 2608.09848
• Published • 6
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
Paper
• 2608.09043
• Published • 7
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper
• 2608.09888
• Published • 596
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Paper
• 2608.07169
• Published • 49
Motif 3: Technical Report
Paper
• 2608.09119
• Published • 43
Scaling Inherently Interpretable Language Models
Paper
• 2608.07594
• Published • 19
Stealing Reasoning Traces from Proprietary LLM APIs
Paper
• 2608.09867
• Published • 99
Evo-Bench: Can Language Models Improve Agent Harness?
Paper
• 2608.09096
• Published • 17