Measuring the Self-Reported Impact of Early-2026 AI on Tech Worker Productivity
A survey of 349 tech workers found median self-reported value gains of 1.4-2x and speed gains of 3x from AI tools. Respondents retrospectively estimated 1.3x value gain for March 2025, 2x for March 2026, and forecast 2.5x for March 2027. The study distinguishes between 'value' and 'speed' gains and cautions that survey results may overstate reality.
Summary
In February–April 2026, we ran a survey of 349 technical workers (including 87 software engineers, 71 researchers, 129 academics and PhD students, and 48 founders and managers) about their usage of AI tools. Compared to previous work, our survey is one of the more detailed surveys of technical workers’ self-reported gains from frontier AI tools.1
We attempt to capture gains due to AI in terms of ‘value’ (how much more value are you creating with AI), rather than ‘speed’ (how long would it have taken you to do these tasks without AI). These can give different answers in principle, in particular if using AI changes the distribution of tasks you work on. For example, researchers could use AI to quickly build an interactive dashboard for their data, which would have taken significantly longer without AI but isn’t that important for their project. We provide more detail on the distinction between value and speed gains in our previous research.
We think that the distinction between ‘value’ and ‘speed’ gains is important because value is closer to the idea that survey designers typically care about, whereas our sense is that it is common for respondents to think in terms of speed, and we expect that speed changes would typically overstate value changes. See methodological details here.
Participants self-reported a median 1.4–2x change in the value in their work due to AI tools. The median self-reported speed change (which we expect to be higher than value change) is 3x.
Using the same question wording, respondents retrospectively estimate 1.3x value of work in March 2025, estimate 2x in March 2026, and forecast 2.5x for March 2027.
Responses are broadly internally consistent – correlations between different value change measures are moderate, only 3% of respondents showed anomalies in their answers that we feel is sufficient to remove them.
There are tentative reasons to be skeptical of the magnitude of responses.
METR staff give the lowest change in value answers of any subgroup we study. We expect that this might be due to METR staff having in mind past findings of gaps between perceived and actual AI-driven productivity.
Early qualitative investigation of public outputs plausibly suggests that especially high change in value estimates are overstated.
See detailed results here.
Importantly, survey results are not necessarily grounded in reality. There are reasons to be skeptical of people’s responses to counterfactual questions such as about AI’s effect on productivity — for instance, our study in early 2025 found that people overestimated AI’s effect on their time spent on tasks by 40 percentage points on average.
More broadly, public estimates of productivity impacts from surveys have tended to be greater than those from field experiments or quasi-experiments. However, it is difficult to determine the extent to which surveys overestimate productivity gains relative to experimental data.2
Nevertheless, we think that surveys provide a useful source of information about the true impact of AI on productivity. Surveys complement other evidence on AI capabilities – benchmarks, RCTs, data from wide deployment – each with different blind spots (more detail here). Surveys are cheap, broad, and straightforward to run inside major AI companies. Surveys can also obtain information about some types of questions that would be otherwise intractable (e.g. qualitative questions).
We would like to see those interested in tracking the potential automation of AI R&D, such as frontier AI companies, run high-quality surveys building on this work. We particularly recommend carefully defining the metric to be estimated, such as by building on the survey questions discussed here. We also suggest prioritizing surveying managers or productivity researchers over individual contributors. See more detailed recommendations here.
How surveys complement other evidence
Different sources of evidence on AI capabilities come with different pros and cons. For example:
Benchmarks are standardized and highly replicable. They’re relatively expensive to make and then somewhat cheap to use once built, and can be cheaply rerun with modifications to understand the impact of methodological choices. On the other hand, there are often concerns about external validity (in particular, that benchmark results are an overestimate of capabilities observed in the wild); it’s challenging to think about how to map estimates to downstream quantities we care about; and they might capture a narrow slice of some larger distribution of tasks (in particular, those that are easier to cheaply evaluate).
RCTs are carefully controlled and perhaps highly externally valid. On the other hand, they too produce an estimate that can be somewhat hard to map to downstream quantities we care about3, and they’re very expensive.4
Observational evidence from broad deployment is very large, relatively cheap to collect, and far-reaching in terms of task distribution. However, it typically suffers from selection effects and is often hard to reason about.
We think of surveys as a flawed but useful complement to other forms of evidence on AI capabilities. As discussed in the summary, surveys have problems – most severely, that respondents have difficulty answering complex quantitative questions accurately – but they are cheap to run and can be used to study some types of questions that would be otherwise intractable.
Methodology
Here is the survey. Take it if you want! The key questions are:
One question conceptualizing increase in speed due to access to AI tools around March 2026.
Three questions conceptualizing increase in value produced due to access to AI tools around March 2026, with estimates for March 2025 and March 2027 requested for one of these.
The full questions details are in the appendices.
The survey additionally asks of respondents:
Their type of work, primary development environment, examples of use and non-use of AI, allocation of work time by activity, and how much the respondent uses AI by activity.
Their duration of experience programming, using LLM-based tools for programming, and using AI agents for programming.
Hypothetical willingness-to-accept for AI tools and actual spending on AI tools.
Other questions to gain a broader understanding of respondents’ work and experience with AI tools.
This survey is larger-scale and more detailed than past surveys of frontier AI tool usage like those used by Anthropic.
One innovation in our survey is attempting to distinguish between ‘value’ and ‘speed’ uplift. Value is defined holistically, in terms of what “you, your team or your leadership would find valuable”. Speed is instead the raw difference in how long tasks take.5 We elaborate on the distinction between value and speed gains in previous research.
Speed measures may differ from value measures by, for example:
Being inflated by individuals doing additional tasks which AI can do well or quickly but which would otherwise not have been worth prioritizing (substitution).
Failing to capture ways in which AI makes higher-value tasks possible.
Not accounting for decreasing or increasing returns to completing tasks.
We think value as defined in this survey is materially closer to what we care about – a multiplier on employee contribution to AI R&D progress. Speed measures are likely biased upwards with respect to value measures; however, value is also more opaque and harder for respondents to think about.67
Another innovation is attempting to triangulate the truth using measures of internal consistency. We use multiple measures and check for consistency.8 However, these checks are not decisive, and, more fundamentally, even when reporting on the internal consistency of perceptions we have little means to test whether perceptions match reality.
Participants come from a convenience sample. They were largely sourced from GitHub, academics via institutional or conference directories, METR, METR staff’s professional networks, and X (formerly twitter). Response rates are typically low – approximately 2% for respondents we email, although very significantly higher among METR staff and their networks – so our results plausibly suffer from significant selection bias.9 Approximately 70% of participants were paid to take the survey (per-hour or fixed payment); of participants who were paid, average pay is $200.
Respondents averaged 12 years programming, 19 months using AI for programming, and 7 months using agentic AI coding tools. 48% are US-based. 50% regularly use Claude Code. Academics and PhD students represent 37% of the sample. Around half of respondents are based in the US and one-quarter in Europe.
In a later section, we display some evidence for higher perceived uplift among people who currently use AI more, are more experienced using AI tools, or work at startups; we see lower perceived uplift among METR employees.
We flag answers that suggest poor data quality (extreme inconsistencies between questions, logical impossibilities, etc.) and filter out 10 respondents who have many flagged responses. See anomalies section for more information.
Results
Productivity gains around March 2026
Key questions
Our 3 measures of value uplift have medians between 1.4x and 2x.10
We are surprised by the ordering of value uplift results. The “fraction of the value” question was intended to be specified so that answers should be approximately equivalent to the reciprocal of the “how long would it have taken you” question, but we see a meaningful discrepancy between answers to each question. On the other hand, we anticipated that answers to the “how many copies” question would naturally be larger than answers to the “how long would it have taken you” question due to costs of parallelizing across people, but we do not find this to obviously be the case in respondents’ answers.
We find it less surprising that answers to the speed measure, which considers the change in time it takes to complete tasks after AI is available, are larger in magnitude. We had hypothesized (following our earlier research differentiating speed and value uplift) that AI users were substituting into lower-value tasks which had become much ‘cheaper’ as a result of AI, causing them to get less value out of AI usage than one might naively expect before accounting for changes in the task distribution.11
Note that different respondents find different versions of the value uplift question most intuitive; results are broadly similar if we use whichever answer they gave to their preferred question.
Validation/consistency checks between measures
We asked value uplift questions in several different ways partly for imperfect consistency checks, and partly because we found early on that, as above, different respondents find different versions of the value uplift question most intuitive.
Note that it’s not necessarily straightforward to interpret differences in answers between questions as a consistency check (although we do largely have this interpretation):
Respondents could anticipate that we are running these checks, which means they might incorporate this fact into their answers.
The questions are asking about slightly different quantities. For example, differences between the questions “How long would it have taken you, in months, to deliver equally valuable work to that which you delivered last month if you had not had access to AI?” and “If your team had to replace you with people just like you (same skills, same knowledge) except that they did not have access to AI, how many copies of you would it need to hire?” might reflect diminishing returns to copies of yourself.
Given the above, we should expect different value measures to be correlated but not perfectly so, which is what we see in the data.
Productivity gains over time
The median respondent retrospectively claims that their uplift was around
[truncated for AI cost control]