Agent TARS is a general multimodal AI agent that brings the power of GUI Agent and Vision into your terminal, computer, browser, and product, providing a workflow closer to human-like task completion through multimodal LLMs and integration with real-world MCP tools. It primarily ships with a CLI and Web UI. The TARS stack also includes UI-TARS Desktop, a desktop application providing a native GUI Agent based on the UI-TARS model.
A general multimodal AI agent for human-like task completion across terminal, computer, browser and product, using GUI Agent and Vision capabilities and seamless integration with real-world MCP tools; primarily used via a CLI and Web UI. UI-TARS Desktop provides local and remote computer and browser operators.
Although the entrypoint code surface showed no autonomy signals, the README clearly describes Agent TARS / UI-TARS as a multimodal GUI Agent stack designed to autonomously operate a computer, browser, and terminal to complete tasks in a human-like workflow. The v0.3.0 release notes explicitly add streaming support for shell command execution and an AIO Agent Sandbox as an isolated all-in-one tools execution environment — indicating the agent runs multi-step actions (shell commands, GUI/browser control) without per-step human confirmation. This is a fully agentic operator stack, not advisory output. No code-enforced approval gate is evident; the sandbox is for isolation, not human approval.