MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval Paper • 2608.30949 • Published 7 days ago • 8
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 5 days ago • 9
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 5 days ago • 521
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models Paper • 2609.00551 • Published 6 days ago • 14
DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory Paper • 2609.00768 • Published 6 days ago • 20
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement Paper • 2609.01481 • Published 6 days ago • 17
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching Paper • 2609.01404 • Published 6 days ago • 26
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 7 days ago • 373
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Paper • 2608.27454 • Published 11 days ago • 30
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published 11 days ago • 82
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published 12 days ago • 177
ReWorld: An Interactive World Model with Long-Horizon Memory Paper • 2608.23565 • Published 14 days ago • 24
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Paper • 2608.24053 • Published 13 days ago • 70
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale Paper • 2608.20634 • Published 17 days ago • 12
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 18 days ago • 274