Technology · · 4 min read
Meta uses code-change data to assess AI’s impact on engineering
A conversation with Meta researcher Moritz Beller examines the company’s developer-productivity measures, the effects of AI and the limits of automated testing.
Meta is using measurements taken from individual code changes to understand how engineering tools and artificial intelligence are affecting developers’ work, according to a report from newsletter.getdx.com. Researcher Moritz Beller says the company has seen the time spent authoring a change fall by more than 40% year over year, even as engineers produce more and larger changes.
The measure, known as diff authoring time, captures the active effort involved in creating, testing and reviewing one code change. Looking at changes individually gives Meta a more focused view than calculating productivity from everything a developer does during a day.
Meta does not use the measure to rank or assess individual engineers. Instead, it examines patterns across teams, departments and the company as a whole. The aim is to spot regressions, identify useful improvements and decide where further investment in engineering tools is justified.
A more detailed view of productivity
The company treats diff authoring time as one signal among several. Throughput, the size of changes and their quality are considered alongside it, because no single measure can describe engineering productivity reliably. A shorter time to complete a change, for example, is not necessarily beneficial if the resulting code is poor or requires substantial correction.
This approach also allows Meta to test whether changes to frameworks and developer tools produce practical benefits. The report cites automatic memoization in the React compiler as an example. Compared with adding caching by hand, the compiler’s automatic approach reduced the time needed to author a change by about 30%. Beller presents results of this scale as evidence that fundamental improvements to the development environment may matter more than minor refinements to interfaces.
The fall in authoring time is not fully explained by faster typing or code generation. Beller believes developers may be using the time they save to gather context and handle tasks surrounding implementation. Prompt writing appears to account for only a small share of that additional activity.
As generating code becomes less expensive, the harder problem may be defining what the code should do. The value of a developer’s intent increases when an automated system can produce an implementation quickly but may misunderstand an incomplete request.
What productivity measures miss
Traditional telemetry is better at recording code changes than at capturing the work that precedes them. Whiteboarding, brainstorming and architectural discussions can be crucial to an outcome without leaving a clear connection to a later diff. Meeting transcripts and other records created by AI could eventually make some of this preparation easier to study.
That does not make architecture less important. Avoiding a design document may simply shift the burden to reviewers, who then have to infer the intended structure from the finished implementation. Faster production can therefore conceal extra work elsewhere in the engineering process.
Beller’s earlier Mind the Gap study compared automatically recorded activity with developers’ own assessments of their productivity. Coding time was a significant predictor, but rest, interruptions and on-call duties also influenced how productive people felt. A new version would need to examine how developers work with agents, how much confidence they place in them and whether coordinating several automated tasks at once creates excessive cognitive strain.
More output also does not guarantee a better working experience. The discussion notes that productivity can increase while concentration, mental effort and overall satisfaction remain unchanged or decline. Interpersonal relationships, shared understanding and informal knowledge exchange are important to engineering work but remain difficult to represent in most metrics.
Reducing unnecessary meetings may improve throughput, yet removing too much collaboration can deprive teams of ideas and context. Small social exchanges matter as well: research discussed in the episode found that informal conversation before and after meetings predicted people’s reported productivity more strongly than internet quality.
The testing challenge for AI agents
AI agents make it cheap to create unit and end-to-end tests, potentially addressing a longstanding gap among developers who previously wrote little automated coverage. But a collection of passing tests can provide misleading reassurance when the implementation and the tests share the same mistaken interpretation of an ambiguous request.
An agent may produce code and tests that are internally consistent while still failing to meet the developer’s actual goal. That risk makes it more important to define correctness before implementation begins. The conversation suggests that specification-driven development could become a counterpart to test-driven development as agents take on more coding work.
The speed of automated implementation also raises the risk that teams unknowingly build the same thing in parallel. Better measurement will need to account not only for code completed, but also for alignment, communication and duplicated effort. Meta’s diff-level data offers a way to see some effects of new tools, but the wider discussion makes clear that engineering productivity remains a human and organisational problem as well as a technical one.