LVMT: Video Mask Transformer for Long-term Video Segmentation Paper • 2609.34895 • Published 12 days ago • 19
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 24 days ago • 43
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published Sep 3 • 182
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models Paper • 2608.13049 • Published Aug 13 • 21