🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Agents · Evaluation

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

First page
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
The curator’s take

Panagiotis Kasnesis and colleagues (University of West Attica) evaluate 9 models from 0.8B parameters to a hosted frontier model at each of the five LLM call sites of Wactorz, a deployed open-source multi-agent home-automation framework. Accepted at a NeurIPS 2026 workshop on SLMs for agentic systems.

Ask this paper

Key points
01

Setup. 280 cases and 2,520 scored calls using the framework's unmodified production prompts and two real Home Assistant installations. Call sites cover intent routing, action classification, device grounding, pipeline planning and code generation.

02

No single ordering. Capability ranks differ by call site and are not monotonic in size: one 4B model does worse than its 2B sibling at grounded actuation.

03

Where hosted wins. The best local model is statistically indistinguishable from both hosted models at four of five sites. Only code generation separates them (p = 0.039 against a small hosted model, p = 0.002 against the frontier one).

04

Safety metric. Aggregate accuracy hides actuation failures: Gemma4 E2B actuates on 87.2% of requests for devices the site does not own, while another model refuses everything.

05

Routing. Routing each site to its best local model scores 91.8% against 95.4% with zero per-call cost. In a live deployment, hosting only the two generative sites matched hosting everything (39/43 each) at 28% of the spend.

Abstract

An agentic system issues several structurally different kinds of LLM calls. It routes intent, classifies actions, grounds language in a device registry, plans multi-agent pipelines and writes the Python code those pipelines run. The difficulty of these call sites varies by an order of magnitude, yet in practice a single model, chosen for the hardest site, serves all of them. In this work, we evaluate 9 models from 0.8B to a frontier hosted model across the five call sites of a deployed open-source home-automation framework (Wactorz), using its unmodified production prompts and two real Home Assistant installations (280 cases, 2520 scored calls). We find that capability is not ordered the same way at every site, and that larger models are not uniformly better: one 4B model is worse than its 2B sibling at grounded actuation. Paired testing shows the best local model to be statistically indistinguishable from both hosted models at four of five sites. Only code generation separates them, against a small hosted model (p = 0.039) as well as a frontier one (p = 0.002). Aggregate accuracy also hides a safety failure specific to actuation, where small models resolve the accuracy/refusal trade-off in degenerate ways: one model (Gemma4 E2B) actuates on 87.2% of requests for devices the site does not own, while another refuses every request it receives. Routing each site to its best local model reaches 91.8% against 95.4% at no per-call cost. In a live deployment judged by a user, hosting only the two generative sites matches hosting everything (39/43 against 39/43) for 28% of the spend, and the actuation gap the benchmark predicted appears as exactly one case in twenty-six. Benchmark, harness and all records are released at this https URL.

Every Monday
Get next week’s papers.
Subscribe on Substack