In partnership with

Last edition covered what people ask AI to do — mostly requests for help, not handovers. This one's the flip side: what AI can actually finish on its own, unsupervised, when someone else is paying and grading. (Edition 282 has the asking-side data, if you want it — no need to reread it here.)

The number: 240 real, already-completed freelance projects, worth a combined $144,000, pulled from a benchmark called the Remote Labor Index (RLI). Real Upwork briefs, real input files, real professionals who already did the work for pay. AI agents were handed the same brief and graded, by humans, against the professional's actual deliverable: good enough that a real client would accept it, or not.

The result: 15.8%. That's the current best score, set by Anthropic's newest model (Claude Fable 5). Eight months ago, when RLI launched, the best any model could do was 2.5%. That's roughly a sixfold jump — real, measured progress on real paid work, not a lab benchmark.

Who's keeping score, and why it matters: RLI is run jointly by the Center for AI Safety (CAIS), a nonprofit with no product to sell, and Scale AI, a company that sells AI evaluation and data-labeling services for a living — including running RLI's own grading. That's not a reason to distrust the number (47 researchers put their names on the underlying paper, and it's public), but it's worth knowing a commercial AI-data company has a direct stake in a benchmark about how close AI is to replacing paid human work.

The methodology, briefly: 358 verified freelancers, 550 candidate projects narrowed to a final 240 across 23 categories — 3D and CAD, architecture, graphic design, video, audio, data analysis, web apps. Every deliverable is graded entirely by trained human evaluators, not an AI judge — the researchers tested an automated judge and found it overstated automation rates by 2 to 3 times, because grading the work turns out to require the same real-software skill the AI workers are being tested on. Evaluators agree with each other 94.4% of the time.

An entire ad agency in the palm of your hand.

Your next campaign needs a dozen fresh ad variations by Friday. Your agency quotes two weeks and a five-figure invoice. Your in-house designers are already buried under this quarter's requests.

Hightouch Ad Studio fixes that. It reads your brand guidelines, your best-performing creative, and your product catalog, then generates on-brand ads your team can ship the same afternoon. You review and approve every asset before it goes live, so quality holds.

Growth teams use it to build variations for every audience, test more of them, and stop rationing creative because production got expensive. The work that once needed a full agency retainer now runs inside your own workflow, at your pace and under your direction.

You direct the work while Ad Studio handles production, and your designers get their week back.

Here's the actually useful part: where the 84% breaks down. It's not random. In Scale AI's own review of failed submissions, 45.6% simply weren't professional quality, 35.7% were incomplete (truncated videos, missing files), 17.6% had technical problems (corrupt or empty files), and 14.8% broke consistency across files in the same project — often more than one of these at once. The newest results show the same pattern persisting even as scores climb: CAIS's own review of Fable 5's redesigned engagement ring calls it "qualitatively much better" than earlier AIs and still unprofessional — a low-effort prong design. On a separate renovation project, GPT-5.5 turned in a bathroom render that looked finished, except it had been faked with an image generator instead of built from the actual 3D model — the kind of shortcut a client only catches by opening the real files.

The pattern that emerges: agents are getting reliably good at generating something plausible from a blank page — first-draft writing, images, audio, code from scratch. They're still unreliable at editing existing work without breaking it, following someone else's precise multi-step brief exactly, and operating unfamiliar professional software the way a real freelancer would. That's the boundary right now — not "AI can't do real work" and not "AI is replacing freelancers" — a specific, checkable line that moved sixfold in eight months and is still nowhere close to the finish line.

If you're deciding whether to hand a real project to an AI agent this month: the safer bet is still "generate me a first draft," where good-enough-to-revise is the bar. The riskier bet is "here's an exact 12-step brief, don't deviate" — that's exactly the kind of work still failing 84% of the time.

We built the full category-by-category breakdown — which kinds of work RLI's agents actually pass, and why — as a live interactive on the site, plus the complete model-by-model leaderboard and a free copy-paste prompt to help you judge whether your own AI's finished work is actually done.

10x the context. Half the time.

Speak your prompts into ChatGPT or Claude and get detailed, paste-ready input that actually gives you useful output. Wispr Flow captures what you'd cut when typing. Free on Mac, Windows, and iPhone.

Keep Reading