Asana Uses GPT-6 Astra in Codex to Cut Browser Agent Costs by 76x

Through StackAI, Asana customers automate work by building no-code workflows that navigate websites, fill forms, and gather information. StackAI CTO Frank Hidalgo, PhD, wanted to make the browser agent behind those workflows cheaper and faster. He used GPT-6 Astra in Codex to map the codebase, run experiments, and compare results; a research effort he estimates would have taken one to two months by hand took about a week. The resulting 144-run study compared GPT-6.1 Sol with three other frontier models, labeled Models A, B, and C. The optimized workflow on GPT-6.1 Sol averaged $0.47 in estimated model costs and about four minutes per run, which Asana reports as 76x cheaper and 5x faster than the original production setup on Model B.

GPT-6 Astra‘s initial investigation found that the agent cached fixed instructions and tool definitions but not the growing history of page text and screenshots it collected, so every request resent that history at full price. It also dropped older screenshots and trimmed text at nearly every step, altering the history so caching alone would not solve the problem and discarding facts that could force the agent to revisit pages. Hidalgo reviewed the proposed fixes and chose three to test: extending caching to the browsing history, increasing the amount of text retained, and removing screenshots in batches instead of at every step.

Because the code was not built for controlled experiments, GPT-6 Astra first refactored it so one frontend and backend could run many workflows in parallel with separate settings. The study covered history budgets of 120,000 and 480,000 characters and six caching and screenshot policies, each tested three times on each of four models, for 144 runs. Every configuration performed the same task: collecting six fields for each of 32 books from a public demo catalog. The best policy let screenshots accumulate to 20 before cutting back to the most recent one, leaving earlier history unchanged for longer stretches; combined with the larger history budget, this became the optimized workflow. All requests, usage records, outputs, and traces were recorded in Command, Asana’s software delivery platform, and findings were turned into tickets and pull requests before production.

For Model B, the optimization reduced estimated model cost from at least $36.21 per run, a lower bound because some original runs hit the step limit before finishing, to $1.24 per run, a 29x reduction. On GPT-6.1 Sol, the optimized workflow was still 2.6x cheaper at $0.47. Every run in the optimized configuration completed the task and returned the correct answer. On GPT-6.1 Sol alone, the larger history budget and new caching/screenshot policy cut cost 4x, from $1.97 to $0.47 per run. Each call was about 3x cheaper because 89% of the input came from cache at 5% of the uncached price. Runtime dropped from at least 22.5 minutes on Model B’s original setup to roughly four minutes. History retention also affected whether the agent produced an answer at all: with the smaller history budget, GPT-6.1 Sol produced an answer in 3 of 18 runs; with the larger budget, all 18 runs produced the correct answer. The article notes that some chart bars and standard-deviation whiskers were approximate reconstructions from the source image.

Asana has shipped the browser-navigation changes in StackAI and is building tooling to make similar experiments repeatable, with plans to fold this kind of testing into the platform’s evaluations so teams can compare cost, runtime, and answer quality when configuring agents. The company is also using GPT-6 Astra in Codex to test product features before release, navigating the platform, trying different inputs, and reporting bugs for human QA reviewers. Hidalgo frames the shift as one from shipping speed to human attention: every engineer becomes a PM leading a fleet of agents.

Asana cuts model costs 76x in browser tests with GPT-6.1 Sol

View Original