All answers

Answers · Updated August 16, 2026

AI productivity statistics: what do controlled studies show?

Controlled studies do not support one universal AI productivity percentage. Reported effects range from 13.8% more customer-support issues resolved per hour and 40% less time on selected writing tasks to a 19% slowdown for experienced open-source developers working in familiar repositories. The defensible conclusion is task-specific: measure completion time, accepted quality, rework, exceptions, and cost on the exact workflow before claiming a gain.

What the main AI productivity studies found

Public figures describe different work. Resolving customer issues per hour is not the same outcome as finishing a writing assignment, merging a software change, or estimating how long a chat would have taken without AI. The percentages below stay attached to their original task, population, comparison, and limitation.

SourcePopulationReported resultMeasured taskImportant limitation
NBER customer support field studyAbout 5,000 support agents+13.8% issues resolved per hourLive customer-support conversationsOne company and one assistive tool; gains differed by experience
MIT professional-writing experiment444 college-educated professionals40% less time; 18% higher rated qualitySelected occupation-specific writingShort assigned tasks, not full jobs or long-running workflows
HBS/BCG consulting experiment758 consultants>25% faster and >40% higher quality inside the AI frontierProduct and strategy assignmentsA separate task outside the frontier produced worse answers
GitHub Copilot controlled experiment95 professional developers55% fasterOne JavaScript HTTP-server taskA single bounded task cannot represent production software work
Microsoft/Accenture developer field experiments4,867 developers across three companies+26.08% completed tasksOrdinary software-development workEach experiment was noisy; the combined estimate is a preprint result
METR experienced-developer experiment16 developers completing 246 tasks19% longer with AIReal issues in familiar open-source repositoriesSmall specialist sample using early-2025 tools
St. Louis Fed workforce surveyU.S. workers ages 18–64Self-reported time savings equal to 1.6% of all work hoursGenerative-AI-assisted work across occupationsSelf-reported counterfactual time, not time-tracked causal measurement
Anthropic Economic Index estimate100,000 Claude.ai conversations81% median estimated task-time savingsTasks represented in sampled conversationsClaude estimated both task duration and savings; off-chat work is missing

Customer support: a measured gain with an experience effect

Brynjolfsson, Li, and Raymond studied the staged introduction of a generative-AI assistant to roughly 5,000 customer-support agents at a Fortune 500 software company. Access to the assistant increased issues resolved per hour by 13.8%. The effect was not evenly distributed: the lowest-skilled and least-experienced workers improved by about 35%, while the most experienced workers saw little benefit and in some cases a small negative effect.

That finding is relevant to assisted support, not a promise for autonomous service. The agents remained responsible for conversations and could accept or ignore the suggestions. Read the NBER working paper. For a business deployment, the comparable measure is not how many replies the model drafts. It is completed, correctly resolved cases per paid hour, with reopened cases, escalations, customer outcomes, and quality review kept in the denominator.

Writing and consulting: faster inside a bounded task

Noy and Zhang randomly assigned 444 college-educated professionals to complete occupation-specific writing tasks with or without ChatGPT. Access reduced completion time by 40% and increased independently rated output quality by 18%. The MIT-hosted study measured short writing assignments such as emails, reports, and analyses. It did not measure every step of a live business process, source verification after the task, or downstream acceptance by a real customer.

A field experiment with 758 Boston Consulting Group consultants adds an important boundary. On tasks inside the tested model’s capabilities, participants completed work more than 25% faster, produced work rated more than 40% higher in quality, and finished more tasks. On a difficult task outside that frontier, AI users were 19 percentage points less likely to reach the correct answer. The HBS working paper calls this uneven boundary the jagged technological frontier.

The operating lesson is not “use AI for all knowledge work.” It is to test representative normal, difficult, ambiguous, and adversarial cases. Faster work is a gain only when the accepted quality threshold stays fixed and the cost of review, correction, and recovery is included.

Software development: why two controlled results can point in opposite directions

GitHub randomly assigned 95 professional developers to build the same JavaScript HTTP server. The Copilot group finished 55% faster on average and completed the task at a somewhat higher rate. This is credible evidence for one bounded coding assignment; the GitHub experiment does not establish the same effect for architecture, debugging an unfamiliar failure, security review, maintenance, or release ownership.

A later set of randomized field experiments at Microsoft, Accenture, and a Fortune 100 company covered 4,867 developers. Combined, access to an AI coding assistant was associated with 26.08% more completed tasks. The authors report that each company-level estimate was noisy and that less-experienced developers adopted the tool more and saw larger gains. See the Microsoft Research preprint.

METR tested a different setting: 16 experienced open-source developers completed 246 real issues in mature repositories they had worked in for years. With early-2025 AI tools available, they took 19% longer. Before the experiment they expected a 24% speedup, and afterward they still believed they had been 20% faster. The METR study shows why perception is not a substitute for timing accepted work.

These results are not a referendum on every coding tool. They describe different developers, tasks, repositories, tools, and periods. A team evaluating AI-assisted development should count accepted pull requests or equivalent delivery units, then include cycle time, review time, defects, rework, incidents, and maintainability. Lines generated or suggestions accepted are activity measures, not business output.

National estimates: useful context, not a company forecast

The Federal Reserve Bank of St. Louis surveyed U.S. workers about generative-AI use and the extra time they believed the same work would require without it. Pooling three 2025 survey waves, reported savings equaled 1.6% of all work hours, including nonusers. A production model translated that response into a possible labor-productivity contribution of up to 1.3% since ChatGPT’s release. The authors clearly label the assumptions: workers may spend saved time on lower-value activity, while business reorganization could create effects the survey does not capture. Review the St. Louis Fed analysis.

Anthropic analyzed 100,000 Claude.ai conversations and used Claude to estimate human completion time and AI-assisted time. The median estimated saving was 81%, and a model of universal adoption implied a 1.8-percentage-point annual labor-productivity effect. Those are model-based estimates, not observed end-to-end time logs. The analysis says it cannot see off-chat validation, later iterations, or whether the output met its real acceptance standard. The limitation belongs beside the number; see the Anthropic research note.

The U.S. Bureau of Labor Statistics does not publish a measure that isolates AI’s contribution. AI-related effects can appear in labor and total-factor productivity, but so can capital investment, process redesign, demand changes, staffing, and other technologies. The BLS productivity Q&A is a useful guardrail against labeling every change in national output per hour as an AI result.

How to measure AI productivity in your business

Start with a workflow, not a tool. A defensible test compares the same completion unit, quality threshold, population, and observation window before and after release. If the workflow changes at the same time, record that change instead of assigning the full difference to AI.

  1. Name the completed unit. Use resolved case, booked eligible appointment, reconciled invoice, accepted report, or merged change—not prompt, draft, call attempt, or model response.
  2. Freeze the quality rule. Define acceptance, required evidence, review authority, correction conditions, and what counts as a failed or unresolved case.
  3. Record time consistently. Separate elapsed time, active human time, wait time, review, rework, provider delay, and recovery. A faster draft can still create a slower workflow.
  4. Include the full cost. Count subscriptions, usage, integration, monitoring, review, correction, incidents, training, and the work required when the provider changes.
  5. Segment the result. Compare new and experienced workers, routine and difficult cases, high- and low-volume periods, and every important exception class.
  6. Connect capacity to an outcome. Hours saved create capacity. They become financial value only when the business uses that capacity, avoids a cost, increases accepted output, or improves an attributable customer result.
  7. Set a release decision. Expand, revise, or stop based on prewritten thresholds. Do not move the goal after seeing a weak result.

Use the AI ROI calculator to keep capacity, cost, net benefit, and payback assumptions visible. The AI readiness assessment checks ownership, data, authority, integration, evaluation, and operations before a build. For implementation scope and pricing components, review the AI implementation cost guide.

A simple example that avoids false savings

Suppose a team processes 400 documents per month. Before AI, accepted completion takes 20 human minutes per document. After release, drafting takes six minutes, review takes five, 10% of records need eight minutes of correction, and monthly operating cost is $900. The time result should use 11.8 minutes per accepted document, not the six-minute draft. The financial result should value the 54.7 hours of released capacity only if the team can use it, then subtract the $900 and any implementation amortization. If error severity rises, the release can fail even when average time falls.

Method, source hierarchy, and open data

Cognautic searched current live results for AI productivity statistics, opened the original paper or publishing institution, and recorded the reported statistic, direction, unit, population, task, period, sample, comparison, method, source URL, and limitation. Randomized and field experiments are listed separately from surveys and model-based estimates. Publisher summaries are used only when they describe their own research and its design.

We do not average the records. We do not convert a reduction in task time into a claim about revenue, apply a coding result to customer support, treat model-estimated time as observed time, or omit negative findings. Values were last checked August 16, 2026. Source owners control their underlying claims and may revise papers, pages, or methods after that date.

Download CSVDownload JSON

The compilation is available under CC BY 4.0. Cite the original publisher for each result and Cognautic for the normalized compilation. Preserve the task, comparison, and limitation when quoting a record. To report a correction, use the contact page; the content standards explain sourcing, updates, disclosures, and corrections.

From a benchmark to a controlled first release

A benchmark can help form a hypothesis; your workflow determines whether it survives. Cognautic’s AI consulting service maps one process, baseline, risk boundary, implementation choice, acceptance test, fixed written scope, and measurement plan. The AI for small business guide shows practical starting points, while business process automation services describes the controlled path from intake to confirmed destination outcome.

People also ask

How much does AI increase productivity?

There is no universal rate. Controlled studies have reported a 13.8% increase in support issues resolved per hour, 40% less time on selected writing tasks, and 26.08% more completed software tasks across three developer experiments. A separate study of experienced open-source developers found a 19% slowdown. The task, worker experience, tool, quality standard, and workflow determine the result.

What is the most reliable AI productivity statistic?

The most useful statistic is the one that matches your workflow and preserves its study design. For customer support, an NBER field study of roughly 5,000 agents found 13.8% more issues resolved per hour. That result should not be applied to sales, accounting, writing, or software development because those tasks have different inputs, review costs, and failure modes.

Does generative AI make knowledge workers faster?

Often on bounded tasks inside the model's capabilities. In a study of 444 college-educated professionals, ChatGPT reduced time on selected writing tasks by 40% and increased independently rated quality by 18%. In a consulting experiment, participants were faster and produced better work on tasks inside the AI frontier, but were 19 percentage points less likely to reach the correct answer on a task outside it.

Does AI make software developers more productive?

The evidence is mixed. A GitHub experiment found 95 professional developers completed one JavaScript task 55% faster with Copilot. Three larger field experiments covering 4,867 developers found 26.08% more completed tasks. METR found 16 experienced open-source developers took 19% longer on 246 real tasks in repositories they already knew. Scope and context matter.

How should a business measure AI productivity?

Define one workflow and compare the same completion unit before and after release. Track elapsed and active time, accepted output, corrections, rework, exceptions, provider failures, human review, operating cost, and the downstream business outcome. Keep changes in volume, staffing, seasonality, and process design visible so the AI is not credited for unrelated improvements.

Can I download and cite the AI productivity dataset?

Yes. Cognautic publishes the normalized records as JSON and CSV under CC BY 4.0. Cite the original source for each study and Cognautic for the compilation. Preserve the population, task, comparison, period, method, and limitation beside each figure; do not average the percentages into a synthetic AI productivity rate.

Rather not DIY?

Want to measure one workflow instead of buying a benchmark?

If you’d rather have someone build this for you, that’s what we do. Start with a free consult — we map your workflows and name the smartest first move. No pitch, no pressure.

Request a free consult