THE WORLD IS NOT STANDING STILL.RSS
BIG CHANGE.

Markdown edition

# AI coding agents raised code output in a firm study, while completed work lagged

> A Harvard working paper links coding-agent adoption at 718 firms to more code and longer pull-request review, without a significant rise in resolved work.

By BIG CHANGE Editorial

Published: 2026-10-10T00:19:01.727Z
Updated: 2026-10-10T00:19:01.727Z
Canonical: https://bigchange.ai/blog/ai-coding-agents-review-output-study

![Two colleagues discuss a software change beside a monitor turned away from view.](https://bigchange.ai/api/media/file/hero-v2.png)
AI-generated conceptual illustration by BIG CHANGE

A [Harvard working paper](https://fion.ac/jellyfish.pdf) offers a useful check on a common productivity claim: more code written with AI should mean more software delivered. Across 718 firms using the Jellyfish engineering analytics platform, Fiona Chen and James Stratton estimate that adoption of coding agents was followed by 30% more lines of code, 20% more commits and 23% more pull requests per active worker. Their measures of completed Jira issues and epics did not show a statistically significant increase. The paper's current version is dated August 4, 2026; its work-event data end in March 2026.

The mismatch matters to engineering leaders because a pull request enters a queue for review, testing and possible revision. In the same study, the elapsed time from a pull request's submission to its merge rose by an estimated 49% after agent adoption. Requests for changes became more common, and comments per pull request increased. Those findings point to heavier review work, though they do not show that every agent-written change is poor or identify a measured defect rate.

## The big change

- **What changed:** In this firm-level study, coding-agent adoption coincided with substantially more coding activity and more demanding pull-request review, without a statistically significant rise in resolved issues or epics.
- **Why it matters:** Lines, commits and pull requests measure work entering the production process. Resolved work is a later measure. A team judging agents by code volume alone could miss the load that arrives in review.
- **What to watch:** Whether firms can increase review capacity and whether later data show gains in completed work. This working paper follows early adoption through March 2026 and cannot settle longer-term effects.

## Assistants and agents produced different estimates

The authors separate *assistants*, which respond with code suggestions as a developer works, from *agents*, which can take a higher-level task and work through multiple steps before a developer approves the result. They measure assistant adoption through GitHub Copilot and Cursor business-license activation. For agents, they combine Claude Code usage data with signals such as bot accounts and tool signatures in commits or pull requests. Some individual use or unintegrated tools may escape those measures. The agent estimate is the additional association around agent adoption, compared with the earlier assistant-adoption period; it is not an estimate for every individual who uses an agent.

The assistant results were smaller: an estimated 12% increase in lines, 9% in commits and 5% in pull requests. Only the commit result was statistically significant in the paper's main estimates. Agent adoption showed significant increases on all three coding-activity measures. A commit records a code update; a pull request submits a change for review. Neither proves that a feature has reached users. The paper uses resolved Jira issues and larger epics as later-stage output measures, based on the tracked workflow in which teams mark work complete after review, testing and deployment. Jira status remains a proxy for delivery, rather than an independent measurement of what users received.

For agents, the estimated increase in resolved issues was 0.12 per worker-month against a baseline of 3.67, with a standard error of 0.17. The result is statistically indistinguishable from zero. Epic completion likewise showed no significant change. The authors report that their confidence interval rules out an issue-completion increase above 12% of the baseline mean over the period studied. That is a narrower finding than saying agents produce no useful software: modest gains remain possible, and issue counts cannot capture every change in value or quality. The authors tested whether issue size shifted using predicted task-length measures and found no evidence of such a shift in their sample.

## More work reached reviewers

The review measures give the output gap a plausible mechanism. After agent adoption, the paper estimates 3.45 additional days from pull-request submission to merge, relative to a baseline of 7.03 days, a 49% increase. This is calendar time in the review process, not a stopwatch measure of a person's active review minutes. The share of pull requests receiving a formal request for changes rose by about 12 percentage points from a 13% baseline, while comments per request rose by 0.58 from a 1.66 baseline, or 35%. The share of workers who reviewed at least one pull request in a month rose by about four percentage points from a 29% baseline, a 14% relative increase. The comparable assistant estimates did not show significant increases in review duration, change requests or comments.

The paper cannot tell from these metadata alone why a reviewer asked for changes. More submissions can strain a fixed review queue; a change in code quality or review standards could also increase scrutiny. The authors found no significant increase in average pull-request size, which weakens one simple explanation for the extra comments. They did not inspect the code content or directly count defects in agent-generated work. Their two-stage production model explains how writing code faster and changing the review required per draft could together limit completed output. The model is an interpretation of the observations, not a separate test that isolates either mechanism.

AI review tools had also spread through the sample: nearly 80% of firms had used one by March 2026. Yet the paper attributes 23.3% of review comments to AI and finds at least one AI comment on 10.8% of pull requests. Adoption of a review tool therefore does not mean that review has become automatic in these firms.

## What the design can establish

The researchers analyze about 300 million work events from January 2021 through March 2026 at 718 consenting Jellyfish client firms, covering 725,938 workers. They compare outcomes before and after firms adopted assistants or agents with outcomes at firms that adopted later or had not yet adopted, using a staggered difference-in-differences design. Firm and calendar-month controls address some stable differences and shared time trends. The inference still depends on how comparable those firms' paths would have been without adoption. Larger firms adopted earlier, and unmeasured changes could have affected both adoption and engineering work. The study is observational, and its estimates should not be read as a randomized test or a forecast for all software teams.

The employment result needs the same restraint. Using LinkedIn-linked total employment and active workers in Jellyfish for engineering employment, the authors could not attribute a significant change to agent adoption during the observed period. That does not show what hiring will do after a longer adjustment or across the wider labor market.

[Earlier BIG CHANGE coverage](https://bigchange.ai/blog/ai-coding-ci-bottleneck-dependable-delivery) examined a separate engineering account of AI coding and dependable delivery. This study adds a view across firms and distinct measures for coding activity, review and resolved work. Its practical lesson is to track those stages together when assessing agents: a faster first draft changes the amount of work waiting downstream, and this paper does not yet show a matching increase in completed issues or projects.

## Sources & further reading

- [Fiona Chen and James Stratton, *Artificial Intelligence in the Firm: Bottlenecks in Software Production*](https://fion.ac/jellyfish.pdf): primary working paper, current version August 4, 2026. Methods, figures and appendices support the reported estimates and their limits; underlying firm-level data are proprietary and aggregated.
- [Ars Technica's October 9 report](https://arstechnica.com/ai/2026/10/ai-coding-agents-generate-more-code-but-not-more-software/): independent contemporaneous account that brought attention to the paper. The numerical and methodological claims above were checked against the paper itself.
- [BIG CHANGE's earlier AI coding and CI article](https://bigchange.ai/blog/ai-coding-ci-bottleneck-dependable-delivery): related coverage of a separate engineering account. It is context for the delivery question, not a continuation of this study.

## Sources

- [Artificial Intelligence in the Firm: Bottlenecks in Software Production](https://fion.ac/jellyfish.pdf) — Primary 78-page Harvard job-market working paper; work events through March 2026, methods and numerical estimates. Proprietary data are aggregated and anonymized.
- [Ars Technica: AI coding agents generate more code, but not more software](https://arstechnica.com/ai/2026/10/ai-coding-agents-generate-more-code-but-not-more-software/) — Independent contemporaneous report; discovery and context, with study claims checked against the paper.
- [BIG CHANGE: AI coding makes dependable delivery the next engineering challenge](https://bigchange.ai/blog/ai-coding-ci-bottleneck-dependable-delivery) — Earlier BIG CHANGE reporting on a separate Linear engineering account of AI coding and dependable delivery.