The cloud-versus-device argument is becoming less useful.
Microsoft’s October 7 Windows announcements point toward a hybrid product instead: the same agent can use local models for some work and cloud models for other work. GitHub Copilot will be able to make that choice automatically, while developers can still select a local model directly.
That arrangement promises lower token costs, better access to local context, and frontier capability when a laptop-sized model runs out of road.
It also creates a product question that infrastructure teams cannot answer alone: what makes a task belong in one place rather than the other?
Routing is a user promise
An automatic router might consider model quality, latency, memory pressure, cached context, cloud availability, and price. A product team has to add concerns that are harder to compress into one score.
Can the prompt leave the device? Can source code? Does a customer require a specific model or endpoint? Is the user offline by choice or merely experiencing a bad connection? Would switching models halfway through a session discard useful state or change the quality of the answer?
Microsoft says Copilot’s Auto orchestration can consider task context and cache state as it moves between local and cloud inference. It also keeps an explicit local-model option for workflows that need direct control. That manual path matters. Routing is not just an optimization when placement affects privacy, cost, compliance, or reproducibility.
A useful placement policy should therefore have at least four inputs:
- minimum capability required for the task
- data and network boundaries that cannot be crossed
- acceptable latency and cost
- state that must survive a route change
The router can stay automatic without becoming mysterious.
Local inference does not make the workflow local
There is an easy marketing shortcut here: if the model runs on the PC, the work stayed on the PC.
Not necessarily.
Microsoft’s technical description makes the distinction explicit. Model selection, inference, and tool execution have separate boundaries. A local model can still call a remote service, reach a network destination, or run a tool with the signed-in user’s permissions.
That is why Microsoft Execution Containers matter alongside local models. MXC lets developers and administrators define which files and network destinations an agent can use, then enforces that policy outside the agent’s control. The model cannot promote itself because it found a convenient shortcut.
Placement and permission are related, but they are not interchangeable. “Run locally” answers where inference happens. It does not answer what the agent can read, change, or transmit.
I previously wrote that hosting location becomes a product feature when customers care about control boundaries. Hybrid AI makes that boundary dynamic. The product may choose a location again for every task, which means the policy has to travel with the work.
Measure the route, not only the answer
A hybrid product can return a good answer while making a poor placement decision.
Product telemetry should distinguish local inference, cloud inference, tool execution, and network access. Then teams can measure route overrides, unexpected cloud fallbacks, cost per completed task, latency at each boundary, state lost during switching, and failures caused by device constraints.
I would also give users a compact explanation when placement matters: “Kept local because this folder is restricted,” or “Used cloud model because the selected task exceeded local memory.” Nobody needs a routing dissertation beside every response. They do need evidence that an important boundary was respected.
Hybrid AI is appealing because users should not have to pick an architecture before asking for help.
The product still has to know what it promised before it picks for them.
Comments
Comments are powered by GitHub Discussions through Giscus.