MintAct trains one model family for phone, desktop and web control

single source· 1 articles · confidence: medium · first seen 2026-09-17 20:00 UTC

What this means for you

Nothing to act on yet: no weights, no licence, no API, and the 48.9 is reported without an evaluation date. If you are weighing computer-use agents, the claim to test later is one modest-size family replacing three specialists.

MintAct, a family of vision-language models at 2B, 4B and 8B parameters, combines three jobs that are usually separate models: finding the element to click on a screen, navigating several steps across mobile, desktop and web apps, and calling visual tools. The authors report 48.9 on OSWorld-Verified and say the models match per-domain specialists at comparable sizes — their claim, not an independent result. Training ran hundreds of environment instances at once under an asynchronous reinforcement-learning loop. No weights, licence, evaluation date or release date appear in the abstract.

Key facts

  • ·MintAct is a family of vision-language models trained at 2B, 4B and 8B parameters. source
  • ·It covers UI grounding, multi-step navigation across mobile, desktop and web, and visual tool use in one family. source
  • ·The authors report 48.9 on OSWorld-Verified, with no evaluation date given. source
  • ·The authors claim state-of-the-art results across a range of benchmarks at comparable model sizes. source
  • ·Training used an asynchronous RL framework hosting hundreds of concurrent instances across per-domain backends. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire