MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Paper • 2606.00793 • Published Jun 8 • 11
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 15 days ago • 339
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning Paper • 2608.27549 • Published 5 days ago • 40
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning Paper • 2608.27549 • Published 5 days ago • 40
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 15 days ago • 339
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 179
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training Paper • 2603.12255 • Published Mar 12 • 91
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints Paper • 2510.14847 • Published Oct 16, 2025 • 57
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion Paper • 2507.02813 • Published Jul 3, 2025 • 60
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion Paper • 2507.02813 • Published Jul 3, 2025 • 60