AI Applications

Claude Code usage data shows domain expertise, not coding skill, drives success

Anthropic analyzed ~400,000 Claude Code sessions and found that domain expertise, not coding skill, is the strongest predictor of success. Experts achieve verified success more than twice as often as novices, and every major occupation succeeds at nearly the same rate as software engineers when using the tool.

Read MoreClaude Code usage data shows domain expertise, not coding skill, drives success

Project Fetch Phase Two: Claude Opus 4.7 Outpaces Humans on Robot Tasks

Anthropic's follow-up to Project Fetch shows Claude Opus 4.7 completing robotics tasks up to 37 times faster than human teams from eight months earlier — without any human assistance. While closed-loop physical control remains a challenge, the rapid general scaling of LLMs is closing the gap between helpful assistant and autonomous physical agent.

Read MoreProject Fetch Phase Two: Claude Opus 4.7 Outpaces Humans on Robot Tasks

Inside GeneBench-Pro: A Genomics Benchmark for Real-World AI Reasoning

GeneBench-Pro is a benchmark that exposes how current AI models struggle with the multi-step, confounding-aware reasoning needed for real genomics tasks like somatic oncology decisions or single-cell eQTL analysis. Each case study provides full datasets and requires models to correct technical artifacts before drawing conclusions.

Read MoreInside GeneBench-Pro: A Genomics Benchmark for Real-World AI Reasoning

Looker 2026: Agentic BI via Governed Semantic Layers and Gemini Reasoning

Google frames its 2026 Gartner MQ Leader recognition around Looker's shift from passive dashboards to agentic BI. The core insight is that autonomous agents need a governed semantic layer (LookML) and deep reasoning (Gemini 3) to act on trusted metrics at scale, as demonstrated by PayPal's 3,000+ user rollout via Looker's Managed MCP.

Read MoreLooker 2026: Agentic BI via Governed Semantic Layers and Gemini Reasoning