🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Agents

AppAgent

First page
AppAgent
Paper summary

Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

Ask this paper

Key points
01

Multimodal control: The agent reads the phone screen (visual input) and issues low-level touch actions (tap, swipe, type), operating apps the way humans do rather than via APIs.

02

Two learning modes: Learns new apps either via autonomous exploration (discovering functionality through self-play) or by observing human demonstrations.

03

Cross-app generality: Demonstrates proficiency across email, social media, shopping, and creative apps, suggesting that multimodal LLMs can generalize across smartphone UIs.

04

Early mobile-agent blueprint: An early example of the on-device multimodal agent pattern that would become a major 2024 deployment theme.

Every Monday
Get next week’s papers.
Subscribe on Substack