Research: Multimodal Agents Approach Human-Level Performance on Web Tasks
A benchmark of realistic web errands shows frontier multimodal agents completing tasks at roughly 70% of human speed and accuracy, up from under 20% a year ago.
Analysis
The gains come from better grounding, longer action histories and improved error recovery. Researchers caution that reliability on safety-sensitive tasks still lags.
Why it matters
Tracking the capability frontier lets enterprises time their pilots to when an agent can actually finish a job end to end, not just attempt it.
Source: AI Frontier
Related use cases
Customer Support Triage and Resolution
An agent classifies incoming tickets, retrieves relevant knowledge and prior resolutions, drafts a response, and resolves simple cases directly while escalating complex ones with a summary.
L1 IT Helpdesk Resolution
An agent handles common L1 requests through guided conversations, executes approved remediation scripts and opens tickets with full context when escalation is needed.
Automated Sales Lead Qualification
An agent enriches inbound leads, scores them against an ideal profile, drafts personalized outreach and books meetings, handing off warm conversations to reps.