DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
Abstract
Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: https://saidgurbuz.github.io/deskforge/
Community
DeskForge is a controllable desktop environment that composes and explores real Linux applications to generate dense supervision for computer-use agents. Each scene varies the apps, their content, window layout, visual style and resolution. Every visible element is labeled with its type, text, occlusion-aware visible region, owning window and interaction properties, and every click is recorded with the screen before and after it.
DeskForge-1M: 1.2M annotated desktop screenshots, 159.7M element annotations and 917K recorded click transitions across 19 applications, 7 visual styles and 7 resolutions.
Results
- Fine-tuning four VLMs (Qwen3.5-4B, Gemma4-E4B, InternVL3.5-8B, UI-R1-3B) on 200K DeskForge-1M grounding examples improves all of them on every held-out desktop condition and on all five external GUI grounding benchmarks. Qwen3.5-4B gains +11.5 points on ScreenSpot-Pro and +10.1 on OSWorld-G. On the held-out desktops, all four also beat every open model we tested (7B–32B).
- With the same Qwen3.6-27B planner and only the action model swapped, Qwen3.5-4B solves 50 instead of 31 of 119 WebArena-Infinity tasks and 15 instead of 3 of 100 OpenApps tasks.
- An RT-DETRv4-L detector trained on the dense labels outperforms the OmniParser v2 and ScreenParse detectors on out-of-domain GroundCUA screens.
🗂️ Dataset: docling-project/DeskForge-1M
🤖 Models: DeskForge collection
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs (2026)
- UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations (2026)
- UI-Venus-2 Technical Report (2026)
- PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence (2026)
- SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2026)
- MintAct: A Unified Visual Agent for Digital Environments (2026)
- Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.02320 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash