Research & Development World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE

After tokenmaxxing backlash, Microsoft moves more agent work onto the desk

By Brian Buntz | October 7, 2026

Microsoft's local MAI-Code-1.1-Flash generates roughly 60 tokens per second with short prompts on Surface Laptop Ultra, versus just under 40 at a 256K-token prompt. The figures come from Microsoft's testing on a synthetic code-generation workload. Image: Microsoft

Microsoft’s local MAI-Code-1.1-Flash generates roughly 60 tokens per second with short prompts on Surface Laptop Ultra, versus just under 40 at a 256K-token prompt. The figures come from Microsoft’s testing on a synthetic code-generation workload. Image: Microsoft

2026 saw the rise, and quick fall, of tokenmaxxing, the practice of enterprise companies treating AI spend volume as a yardstick for productivity. This summer, Microsoft put its divisions on AI token budgets and told engineers that “tokenmaxxing is not what we are optimizing for.”

Now, in its Oct. 7 announcements, it offered customers a way to avoid the same sticker shock by moving more of the agents’ work onto the PC. It imagines always-on mini PCs running agents, and shared Nvidia DGX Stations, desk-side AI workstations, serving 32 or more people at once. The cloud stays available for work that needs it.

Coding is an early test of Microsoft’s proposition. The company is bringing its MAI-Code-1.1-Flash model to local coding workflows, where an agent can write code, run tests and revise its work through repeated model calls. Running those calls locally removes a per-inference charge.

The catch with local inference has been capability: models small enough to fit in a few dozen gigabytes have generally trailed hosted models on serious coding work. Microsoft’s numbers suggest that gap is narrowing.

In Microsoft’s testing, the on-device version of MAI-Code-1.1-Flash resolved 70.8% of tasks on SWE-bench Verified, a benchmark built from GitHub issues, compared with 72.6% for its higher-precision cloud variant. The quantized model occupies about 53GB, with peak memory use reaching 75.5GB at its full 256K-token context. Microsoft compared it against another open model, a quantized version of OpenAI’s gpt-oss-120b, rather than frontier models.

Microsoft’s new Surface Laptop Ultra uses Nvidia RTX Spark hardware and offers up to 128GB of shared memory. Preorders opened Wednesday at a starting price of $2,599, with availability scheduled for Oct. 16.

GitHub’s HydraFusion is designed to route work between local and cloud models within a coding session, with an experimental preview due later this month in the GitHub Copilot app, GitHub Copilot CLI and Visual Studio Code. The intent is for the laptop to handle the repetitive loop of edits and test runs, while larger models remain on call for harder problems. Developers can also select the local model explicitly.

A coding agent left to keep trying also needs limits on what it can touch. Microsoft made its Execution Containers generally available Wednesday, providing policy-driven containment for local agent processes. Entra attribution and Agent 365 and Intune controls for local agents remain forthcoming; microVM isolation is experimental. And local inference does not make an entire session offline: remote tools and separate cloud requests can still send work beyond the PC.

Apple already has a claim on the desk Microsoft wants to occupy. Its M5 Ultra Mac Studio began reaching customers in September, while a version with 512GB of unified memory is promised for late October.

Tell Us What You Think! Cancel reply

You must be logged in to post a comment.

Related Articles Read More >

Starship’s first orbital flight tests key pieces of SpaceX’s million-satellite data center plan 
Under oath, Google confirms three AI agent test escapes as OpenAI, Anthropic and Meta face NYC lawmakers
What AI-assisted engineering taught us about team design 
Big tech turns to debt to fund the AI buildout 
rd newsletter
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest info on technologies, trends, and strategies in Research & Development.

R&D World Digital Issues

Fall 2025 issue

Browse the most current issue of R&D World and back issues in an easy to use high quality format. Clip, share and download with the leading R&D magazine today.

R&D 100 Awards
Research & Development World
  • Subscribe to R&D World Magazine
  • Sign up for R&D World’s newsletter
  • Contact Us
  • About Us
  • Drug Discovery & Development
  • Pharmaceutical Processing
  • Global Funding Forecast

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search R&D World

  • R&D World Home
  • Topics
    • Aerospace
    • Automotive
    • Biotech
    • Careers
    • Chemistry
    • Environment
    • Energy
    • Life Science
    • Material Science
    • R&D Management
    • Physics
  • Technology
    • 3D Printing
    • A.I./Robotics
    • Software
    • Battery Technology
    • Controlled Environments
      • Cleanrooms
      • Graphene
      • Lasers
      • Regulations/Standards
      • Sensors
    • Imaging
    • Nanotechnology
    • Scientific Computing
      • Big Data
      • HPC/Supercomputing
      • Informatics
      • Security
    • Semiconductors
  • R&D Market Pulse
  • R&D 100
    • 2026 R&D 100 Award Winners
    • 2026 Professional Award Winners
    • 2026 Special Recognition Winners
    • R&D 100 Awards Event
    • R&D 100 Submissions
    • Winner Archive
  • Resources
    • Research Reports
    • Digital Issues
    • Educational Assets
    • Subscribe
    • Video
    • Webinars
    • PharmSci360
    • Content submission guidelines for R&D World
  • Global Funding Forecast
  • Top Labs
  • Advertise
  • SUBSCRIBE