UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Paper • 2608.27456 • Published 12 days ago • 112
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Paper • 2608.13558 • Published 26 days ago • 94
HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning Paper • 2607.15255 • Published Jul 16 • 17
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Paper • 2608.13558 • Published 26 days ago • 94
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Paper • 2608.13558 • Published 26 days ago • 94
HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning Paper • 2607.15255 • Published Jul 16 • 17
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding Paper • 2605.14923 • Published May 14
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents Paper • 2506.01520 • Published Jun 2, 2025
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist Paper • 2511.08521 • Published Nov 11, 2025 • 39
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding Paper • 2509.11866 • Published Sep 15, 2025 • 2
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought Paper • 2505.15431 • Published May 21, 2025 • 2
Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent Paper • 2411.02265 • Published Nov 4, 2024 • 26