AppAgent
Free while signed in. Answers cite the passages they came from.

Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.
Multimodal control: The agent reads the phone screen (visual input) and issues low-level touch actions (tap, swipe, type), operating apps the way humans do rather than via APIs.
Two learning modes: Learns new apps either via autonomous exploration (discovering functionality through self-play) or by observing human demonstrations.
Cross-app generality: Demonstrates proficiency across email, social media, shopping, and creative apps, suggesting that multimodal LLMs can generalize across smartphone UIs.
Early mobile-agent blueprint: An early example of the on-device multimodal agent pattern that would become a major 2024 deployment theme.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack