Home
/
Blog
/
Newsela's CTO proves AI transformation demands stronger engineering fundamentals

Newsela's CTO proves AI transformation demands stronger engineering fundamentals

Photo of undefined
Blog_Newsela_s_CTO_proves_AI_transformation_demands_stronger_engineering_fundamentals_2400x1256_dcaadb16b2

Most engineering organizations started their AI push hoping it would make up for gaps in how they plan, review and measure work. Newsela went in with those foundations already built. Dee Wilcox, CTO at Newsela, leads an AI-first transformation across engineering, data and ML. She started it with baseline KPIs in place, a shared definition of team health and a leadership habit of iteration over hype.

On a recent episode of Dev Interrupted, Wilcox walked through what that preparation bought her team, from the metrics that caught hidden bottlenecks to the cost math that convinced her board. Her experience points to a quieter truth about this moment. AI does not replace engineering fundamentals. It exposes them.

A centralized ML team that embeds instead of gatekeeping

Newsela's ML capability did not spring up overnight. It grew out of a data science team that had existed for years. Then one engineer chose to pursue ML and started building product features with LLMs while most of the industry was still learning what the acronym meant. That early work built real competency in using models for product development, complete with evals.

Today that group operates as a centralized team, but not as a gatekeeper. Its engineers embed directly on product engineering squads. When a recommendation service needs enhancement, an ML engineer joins that team to make sure the evals are right, performance is considered and the correct models are in play. Just as often, the right call is to pull back. The embedded expert is empowered to say a problem is deterministic and needs a different approach entirely. These engineers rarely set the standards for every team, but they get pulled in constantly for their judgment on how to frame a problem.

Taking the guardrails off to learn faster

The early phase was cautious by design. Through the spring, Newsela cycled through Copilot, OpenAI's tools and other options, testing rather than committing. Then the pace shifted. Last fall, with a new leader driving the effort across the whole organization, leadership deliberately removed early guardrails to speed up learning across quality engineering, code review and AI-assisted development. The team stopped moving tentatively and applied its own principles of rapid iteration and constant testing to the tooling question itself.

The embedded model carried straight into experimentation. Rather than dictating rigid standards from the center, ML engineers validate evals, model choice and performance inside each squad, and adoption varies accordingly. Some teams made new tools part of their daily work. A couple of others looked at their workflows, concluded a given tool would not fit, and borrowed its principles anyway. The learning cadence is weekly, a continuous loop of talking, tweaking and changing rather than a single rollout decision.

Engineering KPIs anchor every transformation decision

What kept that gamble from being reckless was measurement. Because baseline KPIs already existed, the team could hold engineer sentiment and delivery data side by side. It could ask how people felt while watching how the SDLC actually changed. That discipline predates the AI push by years, and it started with a conversation rather than a dashboard. Wilcox sat down with her engineering team to agree on how they would measure health, starting from a deceptively simple question. How does a team know it is healthy, and what is true when it is?

The answer landed on a small set of core metrics. Dev cycle time captures how work moves and where it stalls. Code review cycle time surfaces waiting and lost focus. Release defects track quality as new technology reshapes the delivery process. Those became the running dashboard for the transformation.

Metrics also solve a leadership problem that has grown sharper in the AI era. Newsela is an edtech company with a strong cost-optimization culture, where every R&D dollar has to justify itself. Concrete delivery health data lets leadership show the board and CEO that the investment is paying off, replacing anecdote with evidence.

Cycle time reveals the hidden costs of AI adoption

The metrics do not just report progress. They diagnose it. Dev cycle time is a direct read on story size and requirement clarity. "If your stories are too big, your dev cycle time goes up," as Wilcox puts it, and ambiguous requirements push it up the same way through extra rounds of rework. Newsela drove that number from four days to three, and now closer to two.

Code review cycle time earns its place for a subtler reason. "Our code review time, to me, is a proxy for someone who's waiting for feedback," she says. The longer an engineer waits, the more they switch context, and the more often they have to start the problem over when the review finally lands. That number sat near one day last year. Right now it is closer to four.

The climb is exactly the kind of hidden cost measurement is built to catch. When guardrails loosened, PR size spiked. Informal norms such as one PR per ticket and no giant changesets had depended on human reviewers to enforce them, and they stopped holding. One team started shipping 1,500-line PRs, and review time went through the roof. Tracking PR size as its own metric turned a vague complaint into something the team could see and correct month over month.

AI code review is still less consistent than human judgment

For all the gains, the tooling has clear limits, and they cluster around confidence. Non-determinism is the sharpest example. "You run the same security review agent on the same repo in different sessions, you're gonna get different output," Wilcox notes. Running locally versus in production shifts the result again. That variability collides with engineering's basic need for high certainty in the places that matter most.

Design work reveals a similar gap. AI-generated output is still recognizable as such, and moving it toward something that feels human-crafted required pulling in Newsela's own design system. Even that remains less mature than it should be. Working with these tools can feel like pairing with a junior developer one moment and something phenomenal the next, and the difference often comes down to context. A human who has worked in a codebase for a while simply forgets less, while agents need constant reminders to validate against every relevant source.

Accountability keeps these limits from compounding. KPIs run across the whole org, but they also break down by team. When code review cycle time or release defects spike on a squad, the conversation goes straight to its engineering managers and tech leads. Gates exist, but a quick "looks good to me" that bypasses them undoes their value. It also leaves someone else to trace what a bot actually did. The active feedback loop is the real safeguard.

Cost per effective PR ties AI spend to quality

Proving business value has meant getting specific about what counts. One North Star metric gaining traction for AI transformations is cost per effective PR. Dan Lines, LinearB's COO and co-founder, describes it as the dollar amount behind every PR that gets merged and delivered. The word "effective" does the heavy lifting. A PR that causes an incident or has to be reworked does not count, so the metric carries a quality factor that raw output never could. It is also a welcome corrective to the return of lines of code as a proxy, a measure that means almost nothing on its own.

Newsela pairs quality-weighted output with on-time roadmap completion, a metric already in place, to show that AI investment translates into predictable delivery rather than just more volume. Being on time more often is its own proof point.

The most persuasive evidence for the CFO and board came from rethinking how software assets get tracked. Long platform investments are notoriously hard to fund, and an 18-month project is often the fastest way to get an initiative killed, since nobody has 18 months. Now a principal engineer or a string of hack days can produce something visible in weeks. Newsela keeps a register of new software assets created that way, and internal tooling has improved sharply as a result. Teams that once shied away from six-month initiatives are far more optimistic about tackling hard problems.

AI rewards what was already strong

That optimism has limits, and they show up exactly where the fundamentals were weakest. Newsela's current bottlenecks live in discovery, planning and design. Those are the areas where requirements and documentation had grown inconsistent, and inconsistency is nearly impossible to automate. AI rewards whatever was already strong and punishes whatever was already weak.

That is what Newsela's head start really bought. The baseline KPIs, the shared definition of health and the habit of weekly iteration did not make the AI push easy. They made its problems visible early enough to fix. Organizations that treat AI as a reason to double down on measurement, consistency and comfort with change will turn experimentation into lasting delivery health. The ones chasing every new model release will keep finding out which fundamentals they skipped.

For more of Dee Wilcox's thinking on AI SDLC experimentation, engineering KPIs and cost per effective PR, listen to the full episode on the Dev Interrupted podcast.

Your next read