MintAct trains one model family for phone, desktop and web control
single source· 1 articles · confidence: medium · first seen 2026-09-17 20:00 UTC
What this means for you
Nothing to act on yet: no weights, no licence, no API, and the 48.9 is reported without an evaluation date. If you are weighing computer-use agents, the claim to test later is one modest-size family replacing three specialists.
MintAct, a family of vision-language models at 2B, 4B and 8B parameters, combines three jobs that are usually separate models: finding the element to click on a screen, navigating several steps across mobile, desktop and web apps, and calling visual tools. The authors report 48.9 on OSWorld-Verified and say the models match per-domain specialists at comparable sizes — their claim, not an independent result. Training ran hundreds of environment instances at once under an asynchronous reinforcement-learning loop. No weights, licence, evaluation date or release date appear in the abstract.
Key facts
- ·MintAct is a family of vision-language models trained at 2B, 4B and 8B parameters. source
- ·It covers UI grounding, multi-step navigation across mobile, desktop and web, and visual tool use in one family. source
- ·The authors report 48.9 on OSWorld-Verified, with no evaluation date given. source
- ·The authors claim state-of-the-art results across a range of benchmarks at comparable model sizes. source
- ·Training used an asynchronous RL framework hosting hundreds of concurrent instances across per-domain backends. source
What the sources say
- Hugging Face Daily Papers (research) — An arXiv abstract setting out model scales, training infrastructure and one reported benchmark score.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersMintAct: A Unified Visual Agent for Digital Environments2026-09-17