Asana slashes Browser Agent Costs 76-Fold via GPT-6.1 Sol and Prompt Cache Optimization

Enterprise work management platform Asana has achieved a 76-fold reduction in per-run model costs for its StackAI browser automation agents. By pairing OpenAI's newly released GPT-6.1 Sol model with append-only prompt caching strategies, Asana also accelerated agent execution speeds fivefold.

Software engineer using a laptop displaying the Asana platform alongside a prompt-cache optimization graphic, illustrating how Asana reduced browser-agent model costs 76-fold using
Asana reduced its browser agent model costs 76-fold by combining prompt-cache optimization with OpenAI's GPT-6.1 Sol.

Architectural Inefficiencies in Agentic Context Management

As autonomous browser agents become central to enterprise workflow platforms, long-running agent execution loops face substantial financial and latency overheads. In an engineering study released by enterprise work management provider Asana, researchers revealed how subtle prompt modification habits inadvertently destroyed API prompt-caching efficiency across multi-step browser interactions.

Asana’s StackAI browser agent was originally configured to prune DOM tree text and purge previous screen captures at every step to stay within a conservative 120,000-character context budget. However, modifying the prompt structure on every step repeatedly invalidated the unchanged prompt-cache prefix required by LLM API providers. As a result, the browser agent was forced to re-process full, fresh input token payloads on almost every step, causing per-run costs to balloon to an average of $36.21 on its baseline production model.

Batch Pruning and Cache-Aligned History Budgets

To restore caching efficiency, Asana engineers redesigned the agent’s context retention policy around append-only history logs and batch pruning. Rather than dropping individual screenshots and trimming text on every tool execution turn, the revised framework accumulates up to 20 screen captures before executing a single, large-scale pruning pass. This batching approach keeps the prompt prefix intact across roughly 19 consecutive API calls, allowing the model provider's cache layer to serve historical conversation state directly.

Simultaneously, Asana expanded the context window budget fourfold from 120,000 to 480,000 characters. Counterintuitively, expanding the history budget dramatically lowered total run costs. Under the smaller budget, models like GPT-6.1 Sol hit step limits or failed tasks due to truncated history; with the expanded budget and cache-aligned pruning, GPT-6.1 Sol achieved a 100 per cent task completion rate while reading 89 per cent of its input tokens directly from cache.

The Financial Calculus of Frontier Model Swaps

The combination of prompt cache architecture optimization and switching the underlying inference engine to OpenAI's GPT-6.1 Sol reduced per-run agent costs from $36.21 down to $0.47—a 76-fold savings—while accelerating task completion times from over 20 minutes down to four minutes. GPT-6.1 Sol's cached input pricing ($0.10 per million tokens) represents a 95 per cent discount over standard input rates, enabling high-frequency agent tool loops to run at a fraction of standard task expenditure.

The study highlights how enterprise software teams must re-evaluate agentic context handling. As frontier models improve at processing expanded token windows, optimizing prompt-cache hit rates yields far greater financial and performance improvements than aggressively stripping context on every API turn.

For further analysis on agentic workflows and foundation model API engineering, explore our report on the OpenAI Developer Model Guide for the GPT-6 Family and view our dedicated updates under Tools & Applications.

Get the next one by email