THE WORLD IS NOT STANDING STILL.RSS
BIG CHANGE.

Markdown edition

# AI use is spreading beyond coding. Its intensity is harder to prove

> Surveys show wider workplace AI use, while experiments find gains and failures in specific tasks. Usage data cannot yet show that other jobs will match coding's intensity.

By BIG CHANGE Editorial

Published: 2026-10-02T04:10:34.217Z
Updated: 2026-10-02T04:10:34.217Z
Canonical: https://bigchange.ai/blog/ai-use-beyond-coding-workplace-evidence

![A seated person works at a laptop while holding a blank sheet of paper at a wooden table.](https://bigchange.ai/api/media/file/ai-work-beyond-coding-hero-v2-1.png)
AI-generated by OpenAI; BIG CHANGE conceptual editorial illustration.

Nearly two in five employed American adults said they had used generative AI for work in the previous week by the second quarter of 2026. Yet AI assisted with 6.3% of work hours across all employed adults, according to a [national survey tracker](https://fredblog.stlouisfed.org/2026/08/does-generative-ai-save-time-at-work/) from researchers at the St. Louis Fed, Vanderbilt and Harvard. The first figure counts people; the second counts their time. Neither says whether the work was good, whether a company earned more, or whether a job disappeared.

Those distinctions matter when asking whether the rest of work will come to use AI as intensely as software developers. Platform records show wider experimentation, including in legal, sales and marketing. Experiments show genuine gains in some noncoding tasks. They also show that access alone can change little, and that a polished answer can be wrong.

## The big change

- **What changed:** AI vendors now report recurring use and agent activity across more business functions, while independent surveys show broader weekly use among US workers.
- **Why it matters:** Use spreads when tools fit existing workflows, but gains depend on task context and checks. Buyers and workers need outcome measures alongside adoption figures.
- **What to watch:** Repeated task use, verified output quality and net returns beyond coding. Current evidence does not establish when other occupations will reach coding's level of use or what happens to employment.

## The numbers measure different things

The [Real-Time Population Survey](https://fredblog.stlouisfed.org/2026/08/does-generative-ai-save-time-at-work/) asks a nationally representative online sample of US adults aged 18 to 64 about work use. Its past-week use rate rose from 28.2% in the third quarter of 2024 to 39.2% in the second quarter of 2026. The estimated share of all work hours assisted rose from 4.1% to 6.3%. Respondents also estimated that AI saved 2.2% of total work hours in the latest quarter. That last number is recalled time saved, not an observed productivity or revenue gain.

OpenAI sees something different: activity within its own enterprise customers. Its [August Enterprise Signals report](https://openai.com/signals/enterprise-data/) says weekly active Codex users grew 108-fold in legal, 41-fold in sales and recruiting, and 26-fold in marketing between February and June 2026, against fivefold growth in engineering. The rates start from undisclosed bases, so they cannot rank how many lawyers and engineers use AI or the share of either occupation's work it performs. Codex produced 64% of combined Codex and ChatGPT output tokens among enterprise customers in June. Longer agent runs produce more tokens; the token share is neither a worker share nor a measure of value. OpenAI sells these products and has an interest in demonstrating their spread.

OpenAI's separate [study of more than 800,000 US ChatGPT messages](https://cdn.openai.com/pdf/work-at-the-frontier-report.pdf), published in July, classified 16.8% of work-related messages as tasks associated with an occupation other than the user's own. Among messages specific enough to map to an occupation, the share was 43.5%. A [September follow-up](https://cdn.openai.com/pdf/work-at-the-frontier-report-202609.pdf) found that some workers returned to such tasks the next month. These are message shares among sampled users, with occupation inferred from business onboarding information. They show what people ask for and revisit, not that they completed the task or changed their job.

Anthropic's [June Economic Index](https://www.anthropic.com/research/economic-index-june-2026-report) also studies activity on the vendor's own services. Its finer sampling captures workday patterns and differentiates chat, Cowork and first-party API use. The companion survey linked usage to roughly 9,700 Claude respondents with at least five sessions. Anthropic warns that those respondents are not representative of the workforce: computer and mathematical workers made up roughly 30% of respondents but about 4% of US employment. Physical occupations were underrepresented. Its earlier [March index](https://www.anthropic.com/research/economic-index-march-2026-report) found computer and mathematical tasks accounted for 35% of Claude.ai conversations in February. That is a share of one product's conversations, not a share of developers using AI, and coding activity was also moving into the API. Vendor traffic can reveal tasks and frequency within a platform; it cannot supply an economy-wide occupational adoption rate by itself.

## Productivity depends on the task and its checks

In a [staggered rollout](https://doi.org/10.1093/qje/qjae044) involving 5,172 customer support agents, access to an AI conversation assistant raised issues resolved per hour by 15% on average. Less experienced and lower-skilled agents improved in both speed and quality; the most experienced and highest-skilled gained a little speed but had small quality declines. Customers became more polite and were less likely to ask for a manager. The study measured one firm's support operation, where resolution and escalation can be tracked. It did not calculate returns for other service desks.

An [experiment across 66 firms and 7,137 knowledge workers](https://www.nber.org/papers/w33795) randomly provided an AI tool within email, meeting and writing software. In the second half of the six-month study, the 80% of treated workers who used the tool spent two fewer hours a week on email and less time working outside regular hours. The researchers did not detect a broader shift in the quantity or mix of tasks from individual access. Three authors worked for Microsoft Research, and Microsoft, which makes the tool, co-organized the experiment and supplied product telemetry. The result concerns the participating firms, not all office work.

A [field experiment with 758 Boston Consulting Group consultants](https://www.hbs.edu/ris/Publication%20Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf), conducted in 2023 and published in March 2026, found faster, better rated work on tasks deliberately chosen to fit GPT-4's strengths. On a different case requiring interpretation of interview evidence, AI users were more likely to give an incorrect answer. BCG collaborated on the study, which tested defined exercises rather than delivered client outcomes.

In [METR's randomized study](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) of 16 experienced open-source developers working on 246 issues in repositories they knew, allowing early-2025 AI tools increased completion time by 19%. Developers nevertheless believed the tools had sped them up. [METR's later attempt](https://metr.org/blog/2026-02-24-uplift-update/) with newer tools suggested gains, but the researchers judged the estimate unreliable because developers unwilling to work without AI opted out and concurrent agent work complicated timing. The studies therefore cannot establish a single productivity effect for coding across tools and years.

Code has unusually accessible feedback: repositories preserve context and changes, while builds, tests and review can expose failures. Even there, a passing test is not the whole product. In customer support, issue resolution and customer reactions provide other feedback. An email draft may save time without changing a firm's output. A legal, financial or medical recommendation needs checks that fit its consequences, and a buyer needs to account for review work as well as tool cost. These are different measurement problems, so use alone cannot establish equivalent intensity or value.

## Work can change before payroll does

In Denmark, researchers linked surveys of workers in 11 exposed occupations and 7,000 workplaces to administrative records. Their [March 2026 revision](https://www.nber.org/papers/w33777) finds employer chatbot initiatives and reported task changes, but no detectable effect on earnings or recorded hours at worker or workplace level two years after ChatGPT's launch; their estimates rule out effects above 2%. That finding is bounded by Denmark, the occupations and the period studied. It cannot establish that later adoption will leave employment unchanged.

For a worker, recurring use may mean a wider task mix or less time in email. For a buyer, the question is whether the whole workflow produces reliable work at an acceptable cost. For the person receiving a support answer or relying on professional advice, the outcome is whether it is correct and accountable. The available evidence supports wider use beyond coding and measurable gains in selected tasks. It does not yet support a timetable for other jobs to resemble software development in depth of use, or a single productivity and employment effect across them.

## Sources & further reading

- [Real-Time Population Survey adoption tracker, explained by the St. Louis Fed](https://fredblog.stlouisfed.org/2026/08/does-generative-ai-save-time-at-work/), August 27, 2026: US past-week use, assisted work hours and self-reported time savings. It does not observe output quality or revenue.
- [OpenAI Enterprise Signals](https://openai.com/signals/enterprise-data/), updated August 12, 2026: usage inside OpenAI's enterprise customer base, with token and weekly-active-user measures. Growth rates lack published starting counts.
- [OpenAI, Work at the Frontier](https://cdn.openai.com/pdf/work-at-the-frontier-report.pdf), July 2026, and its [September follow-up](https://cdn.openai.com/pdf/work-at-the-frontier-report-202609.pdf): classified ChatGPT messages and recurrence among sampled business users, not completed job outcomes.
- [Anthropic Economic Index, Cadences](https://www.anthropic.com/research/economic-index-june-2026-report), June 26, 2026, and [Learning curves](https://www.anthropic.com/research/economic-index-march-2026-report), March 24, 2026: Claude activity, output types and a selected user survey. Their denominators are Anthropic traffic and respondents.
- [Brynjolfsson, Li and Raymond, Generative AI at Work](https://doi.org/10.1093/qje/qjae044), published online February 4, 2025, in the May *Quarterly Journal of Economics*: measured support-agent productivity, worker skill differences and customer responses during a staggered rollout at one firm.
- [Dillon and colleagues, Shifting Work Patterns with Generative AI](https://www.nber.org/papers/w33795), revised November 2025: randomized workplace access, email time and task mix across participating firms. Three authors worked for Microsoft Research, and Microsoft co-organized the experiment.
- [Dell'Acqua and colleagues, Navigating the Jagged Technological Frontier](https://www.hbs.edu/ris/Publication%20Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf), experiment conducted in 2023 and published March 11, 2026: measured speed, rated quality and errors on distinct consulting exercises.
- [METR's developer experiment](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) and [February 2026 update](https://metr.org/blog/2026-02-24-uplift-update/): randomized coding evidence and the later study's selection problems.
- [Humlum and Vestergaard, Still Waters, Rapid Currents](https://www.nber.org/papers/w33777), revised March 2026: Danish worker and workplace surveys linked to administrative hours and earnings.

## Sources

- [Does generative AI save time at work?](https://fredblog.stlouisfed.org/2026/08/does-generative-ai-save-time-at-work/) — Representative US online survey: past-week use, assisted hours and self-reported savings; no verified output or revenue.
- [Enterprise signals](https://openai.com/signals/enterprise-data/) — OpenAI enterprise telemetry: Codex users by function and token mix. Unpublished base counts and token-heavy agents limit inference.
- [Work at the Frontier: How AI is expanding what people do at work](https://cdn.openai.com/pdf/work-at-the-frontier-report.pdf) — OpenAI classification of more than 800,000 work messages; outside-occupation shares are messages, not completed work.
- [Work at the Frontier: How workers are unlocking new ways of working](https://cdn.openai.com/pdf/work-at-the-frontier-report-202609.pdf) — OpenAI selected-user recurrence study; repeat requests do not establish job redesign.
- [Anthropic Economic Index: Cadences](https://www.anthropic.com/research/economic-index-june-2026-report) — Claude use patterns and nonrepresentative survey of about 9,700 selected users.
- [Anthropic Economic Index: Learning curves](https://www.anthropic.com/research/economic-index-march-2026-report) — February 2026 Claude.ai and API samples. Coding-related 35% is a Claude.ai conversation share.
- [Generative AI at Work](https://doi.org/10.1093/qje/qjae044) — Published QJE study (May 2025 issue) of 5,172 agents in one firm's support operation: 15% more resolutions per hour on average. Gains concentrated among novices; top-skilled workers had small quality declines. Customers became more polite and less likely to request a manager. No general ROI or employment effect.
- [Shifting Work Patterns with Generative AI](https://www.nber.org/papers/w33795) — NBER May 2025 issue, revised November 2025 (day not stated). Randomized study of 7,137 workers in 66 firms. Three authors worked for Microsoft Research; Microsoft made the tool, co-organized the study and supplied telemetry.
- [Navigating the Jagged Technological Frontier](https://www.hbs.edu/ris/Publication%20Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf) — BCG-linked controlled consultant exercises conducted in 2023, published in Organization Science March 11, 2026; task-dependent gains and errors, not client outcomes.
- [METR early-2025 developer productivity study](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) — Randomized 16-developer, 246-issue study; 19% longer with dated tools in a narrow sample.
- [METR developer productivity experiment update](https://metr.org/blog/2026-02-24-uplift-update/) — Researchers explain why later speedup estimates are unreliable due to selection and timing.
- [Still Waters, Rapid Currents](https://www.nber.org/papers/w33777) — NBER May 2025 issue, revised March 2026 (day not stated). Danish survey and administrative study finds early bounded null effects on earnings and recorded hours.
The BIG CHANGE newsletter

The big picture. At your pace.

Recent stories on AI and robotics, the shifts worth watching and practical ideas to use. Choose a daily briefing, weekly digest or monthly perspective.