Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 8 days ago • 88
Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations Paper • 2608.01628 • Published 8 days ago • 22
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 11 days ago • 38
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Paper • 2607.24957 • Published 15 days ago • 19
Self Gradient Forcing: Native Long Video Extrapolation Paper • 2607.20368 • Published 20 days ago • 35
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 140
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 87
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Paper • 2607.05373 • Published Jul 6 • 66
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation Paper • 2606.26087 • Published Jun 24 • 35
World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible Paper • 2606.13652 • Published Jun 11 • 16