H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published 1 day ago • 29
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Paper • 2608.31106 • Published 2 days ago • 93
4DAnyone: Create Anyone in 4D from a Casual Monocular Video Paper • 2608.20335 • Published 13 days ago • 83
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning Paper • 2608.09926 • Published 23 days ago • 14
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published 20 days ago • 46
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models Paper • 2608.10708 • Published 22 days ago • 15
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection Paper • 2608.11562 • Published 21 days ago • 10
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation Paper • 2608.13489 • Published 20 days ago • 98
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published 27 days ago • 59
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Paper • 2608.07468 • Published 26 days ago • 107
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published Jul 31 • 41
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis Paper • 2608.02437 • Published about 1 month ago • 72