Consulting Practice John Epperson Consulting Practice John Epperson

Ten Strategies for Saying No

Ten specific tactics for saying no without burning bridges, and one uncomfortable truth about why we don't say it in the first place.

The reason "no" is a hard word for most people isn't laziness or selfishness. It's that saying no to something almost always means agreeing to a small, immediate cost (a moment of social friction, a moment of feeling unhelpful) to avoid a larger, delayed cost (weeks of energy drain, a project of your own that doesn't get done). The small cost is right in front of you. The larger cost is theoretical. Human wiring goes for the immediate.

But the wiring is wrong. Anyone who has the authority to say yes or no is demonstrating leadership every time they use it. Saying yes to everything isn't kindness. It's a slow way of ceding your priorities to whoever asks last.

Here are ten specific tactics for saying no that I've collected over the years. Some are for one-off requests. Some are for building a culture where the request stops coming. All of them work better than "no, sorry."

1. Anchor to defined goals

The strongest form of no is the one where you can point at what you're doing instead. "I'm working on X, which is our priority this quarter. Adding this would mean stopping X. Do you want me to stop X?" This isn't a rhetorical trick. It's you handing the requester the actual decision they're making, which is a resource allocation, not a favor.

A team with clear vision and goals has this as a default posture. A team without clear goals ends up saying yes to everything because there's no principled reason to say no to anything.

2. Delay it

"Yes, we can do this, after Q3" is often functionally equivalent to no, but it doesn't feel like no to the requester. Sometimes the item genuinely can wait. Sometimes the requester loses interest before the delay lapses. Both outcomes are wins.

The failure mode: don't use delay as a lazy no. If you know you're never going to do it, say so. Fake delays erode trust over time.

3. Buy time to think

"Give me a day to think about it." This is one of the most underused tactics. Most requests come with an implicit "answer now" pressure that isn't real. When you take time to think, you almost always come back with a better answer, and often the request gets refined in the meantime.

The trick is to actually use the time. Otherwise the delay is a stall.

4. Not yet

"I can help with this once you've done X first." Use this when the request depends on a prerequisite the requester hasn't handled. Common cases: you're being asked to build something that needs a design decision from someone else, or you're being asked to fix a bug that hasn't been reproduced.

This is a productive no. It puts the ball back in the requester's court in a way that either gets you what you need to move forward or reveals that the request wasn't as urgent as it seemed.

5. Delegate

"I can't do this, but here's who can." Delegation is a specific kind of no that says "the request has merit, but not with me." Effective when there's genuinely a better resource. Less effective when everyone knows you're pushing it off onto someone who's also overloaded.

6. Create policies

"Requests for this go through the ticketing system" or "That belongs to the ops team." Policies work well when the same request has come to you repeatedly. Instead of saying no ten times, you say no once, structurally, by naming the correct process.

The failure mode: making up policies on the fly to avoid the immediate ask. If the policy isn't real and doesn't get documented, you're just deflecting.

7. Help me say yes

Ask the requester to justify it. "Walk me through why this needs to happen this quarter." "How does this move our top-three goals forward?" You're not being combative. You're making them do the work you'd otherwise have to do to figure out whether to agree. Sometimes they discover the answer is "it doesn't really" and withdraw the request. Sometimes they build the case and you agree because now you understand it.

8. Appeal to budget

"I'd love to do this. Get me the budget for a contractor and we're in." Budget appeals shift the request into a resource conversation, which is the correct framing anyway. Requests that don't survive a "who's paying" question often shouldn't have been requests in the first place.

9. Rip the bandage off

Sometimes the right answer is just "no." Clear. Direct. With a brief explanation of why. Nothing padded. No delay. No maybe.

This is harder than the softer versions, but it's often more respectful to the requester, because it doesn't waste their time. And it protects your own energy from the drain of a maybe that never resolves.

The key: always explain why. "No" without reasons feels arbitrary. "No, because we're already committed to X and adding this would delay it" is a decision the requester can accept and adjust to.

10. Say no together

For asks that are systemic (unreasonable expectations, cross-team demands, ongoing pressure), a chorus is louder than a solo. Get the support of your peers, your team, your manager. Present a collective no. This is especially useful when the requester has more organizational power than any one respondent.

Working with people who won't say no

If you manage or work closely with someone who's highly agreeable and who says yes to everything, you have extra work to do. Left alone, they'll take on more than they can carry and burn out. Some things that help:

  • Ask what they're already doing before you add anything. Make them tell you the full list, not just the piece you're adding.
  • Give them permission to say no by saying it for them first. "I don't think this should go to you this quarter. Push it back."
  • Model saying no yourself, especially to their asks. It teaches them that no doesn't damage the relationship.

The other side: what you should say yes to

The point of getting good at no isn't to become a curmudgeon. It's to protect capacity for the things that are actually worth doing.

Say yes to work that moves your goals forward. Say yes to requests that build the relationships you want. Say yes to the small favors that cost nothing and matter a lot to someone. Say yes to the strategic bet even when you can't fully see how it pays off.

Goals that don't change your present actions are weak goals. If saying no to something today wouldn't be justified by the goals you've written down, either your no is wrong or your goals are.


If you're leading a team and feeling like your ability to prioritize is drowning under other people's requests, that's often a solvable problem. It's the kind of conversation Rock Agile has with client engineering leaders as part of what we do. If it'd help to talk it through, get in touch.

Read More
AI + Engineering John Epperson AI + Engineering John Epperson

Why Letting AI Drive Produces Worse Code

The productive AI-generated code and the problematic AI-generated code look identical at first read. That's the whole problem.

The productive AI-generated code I've seen in the last year is genuinely impressive. It compiles. It's tested. It's often stylistically clean. If you were reading it in isolation, you'd think it was written by a competent developer.

The problematic AI-generated code I've seen in the same year has the same properties. It compiles. It's tested. It's stylistically clean. That's the whole problem.

Bad AI-generated code doesn't look bad. It looks fine. It fails in specific ways that don't trip the alarms most teams use to catch bad code, and the failures compound over time in ways that get expensive to unwind. This piece is about what those failure modes are, why they happen, and what senior judgment is actually doing when it's in the loop.

The three failure modes

Almost every bad AI-generated code decision I've watched fits into one of three shapes.

Speculative abstraction. The AI extracts a class, an interface, a helper, a domain model, "in case it gets reused." No second caller exists. The extraction is defended on the grounds of clean architecture principles. What actually results is a codebase with more classes than callers, more indirection than the domain warrants, and more surface area to maintain than the actual complexity requires. Domain models without callers are debt. AI is systematically biased to create them.

Concretely: I recently watched an AI first-pass extraction propose a new domain model based on the argument that "four generation services duplicate this query pattern." On re-grep, only one of the four services was actually doing the query the extraction described. The other three were doing a related-but-different pattern that happened to share a keyword. Conflating them into one domain model would have made the extracted class a god-object serving four different concerns. The pattern-matching that suggested the extraction was superficial. Verifying the claim required manual inspection the AI didn't do on its own.

First-good-enough acceptance. The AI produces a solution that works. Tests pass. It goes in. The problem is that the solution is the middle-of-the-distribution answer, not the right one. Passing tests is not the same as being well-shaped. Passing tests just means the behavior is preserved, which is a lower bar than "the code is designed well."

Concretely again: on the same refactor, the AI's first-pass implementation extracted eight rule classes with an #items method each, and every method rebuilt the item structure inline. Tests passed. Lint clean. But every keyword argument being passed to the item constructor was derivable from self, a textbook data-clump smell. The pattern the code was asking for was Template Method: base class owns construction, subclasses declare properties via hook methods. The AI's first pass didn't see it. It saw "extract methods; make each a class." Which is right, structurally, but stops one layer short of the pattern the code actually wanted. Catching that required me to read the extracted classes again after they were "done" and push back on what the AI thought was a completed refactor.

Pattern-matched architecture over local context. The AI recognizes a familiar problem shape and applies a familiar solution, even when the local context argues for something simpler. "This looks like a service that needs a strategy pattern," even when the actual situation has one strategy and no realistic path to a second. The pattern-matching is wide. It's also blind to why patterns exist in the first place, which is to solve real problems, not to be applied preemptively.

These failure modes have the same root cause. The AI is trained on a corpus of code. That corpus rewards clean-looking abstractions, familiar patterns, and comprehensive coverage. The AI reproduces those defaults. Whether they fit the specific problem in front of you is not a question the AI can answer on its own.

What the empirical data says

GitClear ran a study analyzing about 211 million lines of code change data from 2020 to 2024, comparing AI-heavy codebases against pre-AI baselines. The findings match what I've been watching happen up close:

  • Code churn (revisions within two weeks of the initial commit) grew from 3.1% in 2020 to 5.7% in 2024, with AI-heavy projects seeing about 39% higher churn.
  • Duplicated code blocks rose roughly eightfold in 2024.
  • AI-authored PRs carried about 1.7x more issues per PR (roughly 10.8 versus 6.4) than non-AI PRs in the same codebases.
  • Estimated tech debt increased somewhere in the 30-41% range post-adoption.

None of these numbers say AI is bad. They say unsupervised AI amplifies whatever pattern is already there, and the default pattern for most teams is subtle wrongness. More code, faster, more of it slightly off, more of it copy-pasted rather than reused. Absent a discipline that catches the failure modes, the AI's throughput becomes the codebase's drift rate.

The full GitClear report is public if you want the methodology and the raw numbers. Their "AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones" writeup is the primary document.

What senior intuition is doing

The person in the loop who catches these failure modes is doing something specific, and it's worth naming, because most conversations about "code review of AI output" gloss right past it.

The senior isn't checking that the code is correct. Tests already check that. The senior is checking that the code is shaped right for what it needs to be. That's a different skill. It requires having enough exposure to codebases-in-decay to recognize the shape of decay early. It requires being able to look at a proposed abstraction and say "this hasn't earned its place yet." It requires being able to look at a first-good-enough solution and say "the shape is still wrong, even though the tests pass."

That's not code review. That's design judgment. It's the thing the AI doesn't have.

The most concrete version of what senior judgment is doing:

  • Recognizing when a proposed abstraction has no second caller and pushing back.
  • Recognizing when a solution passes tests but leaves a smell (parameter passing where instance methods would do, imperative construction where declarative would read better, names that leak presentation into the domain layer) and iterating.
  • Recognizing when a pattern is being applied for pattern's sake instead of for the situation's actual needs.
  • Refusing to conflate adjacent-but-different concerns just because the AI's pattern-matching links them.
  • Distinguishing "this could theoretically get reused" from "this is already duplicated in the code AND has a real second caller coming."

Every one of those is a small judgment call. Each one is worth a small amount, individually. Miss enough of them and the codebase drifts. Because the AI's throughput is so high, the drift can happen fast.

Being senior doesn't eliminate this. It reduces it. Even careful, experienced reviewers let things through, some of the time the reviewer notices a smell and rationalizes their way past it because fixing it would slow them down, some of the time the smell is small enough to miss on a busy read. The senior instinct is calibrated on the large-shape stuff, god-object services, misplaced abstraction layers, wrong-named domain models. It's less reliable on the small-shape stuff, a slightly-too-long method, a variable name doing too much semantic work, a curve-fitted test. The small stuff accumulates whether the senior catches it or not. But the reduction rate, even a modest reduction rate, compounds over months and years, and that compounding is the whole difference between a codebase that stays healthy and one that gets rewritten in two years.

The mid-level failure mode, concretely

The place I've watched this go pathological is with mid-level engineers who use AI heavily and don't yet have the senior intuition to catch the wrong shape early. On a recent project, I watched a capable engineer go in circles for weeks. Their pattern:

  1. Hand a problem to the AI.
  2. Get a solution.
  3. Apply it. Tests pass.
  4. Ship.
  5. A day later, an edge-case bug surfaces from the subtle design flaw the AI's solution embedded.
  6. Hand the bug to the AI.
  7. Get a patch. Tests pass.
  8. Apply. Ship.
  9. A day later, a different subtle problem surfaces from the patch.
  10. Repeat.

On paper, this engineer was productive. In practice, they were writing code that needed rewriting almost as fast as they wrote it. The 39% churn number GitClear measured has a concrete shape to it, and this engineer's commit history was exactly that shape.

The engineer wasn't careless. They were doing what the AI encouraged them to do: accept the first solution, ship it, move on. They didn't have the calibrated intuition that would have said "wait, this solution has a shape problem that will bite us in production," so they never asked for a second option. Every solution looked fine when they applied it. The subtle wrongness only revealed itself after the fact.

A senior looking at the same starting problems would have caught the flaws before the first solutions shipped. Or would have pushed back on the AI's first pass and asked for alternatives. The mid-level engineer didn't know to do either, and the AI didn't do it for them.

This is the pattern managers should be looking for. It doesn't show up in "is code shipping" metrics. It shows up in "how often is code shipping that needs to be revised within a week" metrics. If that number is climbing, you have a driver problem, not a productivity problem.

The team-level failure

The most common failure I see in teams that have adopted AI heavily: they haven't articulated who's supposed to be doing the senior judgment step, and by default no one is.

A junior developer using AI is likely to accept the first-pass output. Nothing in their experience trains them to spot the "shape is wrong" smell yet. That's fine, if they're being paired with a senior. It's a problem if they aren't.

A mid-level developer using AI often produces code that reads better than they used to write solo, which reads to management like a productivity win. But if you look at the shape of what's landing, the same three failure modes are in the diffs. The mid-level developer doesn't catch them because their own calibration isn't yet at the level where they would.

A senior developer using AI can produce excellent work, but only if they're using the workflow that requires them to critique the AI's output. If they're using the AI passively, accepting completions, taking first-pass outputs to production, they've effectively demoted themselves. Their senior judgment doesn't help if it isn't in the loop.

The team-level pattern is that AI adoption is often unstructured. Nobody says explicitly "AI output requires a senior review." Nobody says "AI output should always come with the alternatives the AI considered." Nobody says "if the AI's first pass looks fine, that's your signal to re-read it, not your signal to ship." So the drift happens.

What managing this actually looks like

The intervention that catches AI drift is human review at the "is this well-shaped" level. That's not the same as code review at the "does this work" level, and most existing review processes are calibrated to the second, not the first.

Concrete things that seem to help:

Require AI-generated code to come with the alternatives that were considered. If a developer used AI to write a solution, ask them to also show the two other solutions they didn't pick and why. This is the discipline the AI needs to be pushed toward anyway, and it produces evidence that the design decision was made deliberately, not by accepting the first thing that came out.

Explicitly name "shape" as a review dimension. Reviewers should feel empowered to say "this works but the shape is wrong" and have that be a valid review comment. Right now, in a lot of teams, that comment reads as bikeshedding. It shouldn't. The shape of the code is the thing that will still be right (or wrong) in three years.

Watch for the drift signals. Rising abstraction count without corresponding domain complexity is a signal. New classes that only have one caller are a signal. Solutions that pattern-match to famous designs without local justification are a signal. Any of these individually is fine. All of them accumulating is a codebase headed toward a rewrite in two years.

Track code churn as a leading indicator. If your team's commit history shows a rising rate of revisions within a week of initial commit, that's the GitClear number reflected in your own data. Watch it monthly. If it's climbing while AI adoption is climbing, the two are almost certainly connected.

Give seniors permission to slow down AI-heavy juniors. If a junior is producing three PRs a day of AI-generated code, and a senior can only meaningfully review one, the answer is not "let two ship unreviewed." The answer is that the junior needs to slow down. Faster production of unreviewed code is negative value, not positive value.

Make pair-programming an explicit senior/mid-level pairing pattern. Not every mid-level engineer needs to sit with a senior all the time, but for the specific work that requires the senior instinct to shape (a new domain, an unfamiliar codebase, a refactor of load-bearing code), the pairing is what develops the intuition. The AI is not a substitute for that pairing. It is a substitute for typing.

Explicitly reward "the refactor after the refactor." When a mid-level engineer's AI-assisted work needs a senior to come back and reshape it, that reshape is high-value work. Team norms should treat it that way, not as "someone else fixing your work." Otherwise mid-level engineers hide the reshape need, and the drift accumulates unnoticed.

The metrics to watch

If you're a manager trying to figure out whether AI adoption at your company is going in the healthy direction or the drift direction, the metrics that are actually diagnostic:

Code churn rate (commits revised within two weeks of first commit). If it's climbing while AI usage climbs, you have a driver problem. The GitClear numbers give you a benchmark: pre-AI, this was around 3-4%. Above 5-6% is concerning.

PR review comments per PR. If they're rising, your reviewers are catching things, good. If they're stable while the AI-authored PR count is rising, your reviewers are letting more through.

Time-from-first-commit-to-final-shape. Not time to ship. Time to the version of the code that will still be there in six months. If code is shipping fast but getting reshaped repeatedly, the reshaping is where the real cost is.

Duplicated code blocks per feature. If features are adding more duplication than they're removing, the AI's default of regenerating rather than reusing is winning. This is fixable, but only if you're measuring it.

New abstraction count vs. domain complexity growth. If your app is adding domain models faster than it's adding domain, the abstractions are speculative. Prune them or the codebase becomes a maze.

These aren't hard metrics to compute. They're just metrics most teams don't currently look at. The teams that adopt AI without watching them are the teams whose codebases quietly rot.

Positions that aren't defensible

"AI code needs less review because it passes tests" is the position that produces the drift. Passing tests is a floor, not a ceiling. Refusing to raise the ceiling because tests pass is how codebases quietly rot.

"Senior engineers shouldn't have to review AI output" is also indefensible. That's what the senior judgment is for. Skipping that step doesn't save time. It shifts the time to future engineers who have to unwind the drift.

"We can't tell if code is AI-generated so we can't review it differently" is technically true but misses the point. The review discipline should apply to all code. The AI output happens to fail the discipline more often, but that's the review's job to catch, not to know the origin.

"Our team is small and can't afford senior review on every PR" is a real constraint, but the answer isn't "skip review." The answer is to accept that the codebase will drift proportionally to the review skipped, and to plan for the reshaping work that will need to happen later. Every skipped review is deferred debt, not saved time.

The bigger point

AI is a real amplifier. That's not in doubt. The question is what it amplifies.

If it's amplifying the throughput of a senior who's making design decisions well, you get transformative productivity. If it's amplifying the throughput of unreviewed decisions, you get transformative drift. The tool is the same. The presence or absence of judgment in the loop is the whole thing.

Most teams adopting AI right now are optimizing for throughput without asking about the shape of what's being produced. That's a bet. If it works out, they save a lot of engineering hours. If it doesn't, they end up with codebases that look like they were built by twelve people over three years, when in fact they were built by three people over six months. That's a specific kind of technical debt, and the interest rate on it isn't obvious until much later.

The intervention isn't hard. It's just easy to skip.


If your team is figuring out what "good AI adoption" actually looks like in practice, that's a conversation Rock Agile has become deep in. If it'd help to compare notes, get in touch.

Read More
Ruby & Rails Craft John Epperson Ruby & Rails Craft John Epperson

Sprockets and Import Maps, Side by Side: A Practical Coexistence for Legacy Rails Apps

You don't have to rewrite all your legacy JavaScript to modernize a Rails app. Here's how Sprockets and Import Maps coexist without fighting.

If you inherited or maintain a Rails app that predates Rails 7's JavaScript overhaul, you're probably sitting on a pile of legacy JavaScript in app/assets/javascripts/. Custom code. jQuery plugins. Maybe some CoffeeScript. Maybe even Angular. It works. It ships. Nobody wants to rewrite it, and honestly, nobody should have to. It's doing its job.

But you probably also want to add something modern. Stimulus, Turbo, Sentry via CDN, a new library that only ships as ES modules. The official Rails 7 answer for that is import maps. Which sounds great, until you realize the tutorials all assume you're starting fresh, and the upgrade guides mostly assume you're prepared to rewrite everything downstream.

Here's the practical answer for legacy Rails apps: you can run both pipelines side by side. Sprockets serves the legacy bundle. Import maps serve the modern code. If you get the load order right, they don't fight each other, and you get to modernize the parts you want to modernize without touching the parts you don't.

This isn't a hypothetical. It's how Rock Agile actually moves long-lived Rails apps forward. And it started with one specific client and one specific problem.

The specific problem this pattern solved

We inherited a Rails ERP that had been in continuous production since the Rails 2 era. It had been through every major Rails upgrade since: 2 to 3, 3 to 4, 4 to 5, 5 to 6, and eventually 6 to 7. Every one of those upgrades did the right thing for the Ruby side of the codebase. The JavaScript side accumulated. Nothing ever got seriously rewritten, because it worked, because the client didn't want to pay for a JS overhaul that solved no business problem, and because the JavaScript test coverage was thin enough that a rewrite would have been a real risk.

By the time we hit the Rails 7 window, the JavaScript pile in app/assets/javascripts/ was:

  • 51 hand-written .js files in the app tree.
  • 13 CoffeeScript files still compiling through Sprockets.
  • About 6,900 lines of application JavaScript total, spread across those files.
  • Another 2,240 lines of vendored JavaScript in vendor/assets/javascripts/, all of it depending on the Sprockets pipeline to concatenate and serve it.

The vendored libraries alone were a museum of the last decade of front-end Rails: jQuery, jQuery UI, jQuery UJS, jQuery Datepicker, jQuery Timepicker, jQuery TreeTable, DataTables, an old Angular router, Bootstrap 5, Lodash, and Toastr. Some of those are still maintained. Some are barely maintained. A few haven't been touched in years. All of them are load-bearing for at least one screen in the app.

When Rails 6 shipped Webpacker as the default JavaScript pipeline, the community consensus was: rewrite your legacy JS to modules, migrate through Webpacker, and by the way you'll need to add real JavaScript tests along the way. For a fresh app, that's the right advice. For a Rails 7 upgrade of a nine-year-old ERP where JavaScript tests are thin and the client's willingness to fund a JS rewrite is zero, that advice would have added six months of work and material regression risk to what should have been a clean version upgrade.

So we didn't move to Webpacker. We waited. When Rails 7 shipped import maps as an alternative, and later when propshaft became a viable Sprockets replacement for asset serving, a different path opened up: keep the legacy JavaScript exactly where it was, add import maps for anything new, and migrate the app's JS piece by piece as we touched each part for other reasons. No rewrite. No big-bang migration. No JavaScript-testing effort we couldn't sell.

That's the pattern this post is about. It worked so well that we've since used it on a second legacy client with a similar profile, and it's become part of how Rock Agile approaches legacy Rails work generally.

The core pattern

In your layout:

<%= javascript_importmap_tags %>
<%= javascript_include_tag 'sprockets', defer: true %>

That's most of the trick.

The import map tags emit first. They include the modern JavaScript that lives in app/javascript/, imported through config/importmap.rb. Because these are <script type="module"> tags, they run synchronously in module context. Anything they set on window is set immediately.

The Sprockets bundle comes second, with defer: true. That means the browser downloads it in parallel but waits to execute until the DOM is parsed. By the time it runs, the module-context imports have already finished setting up whatever globals the legacy code depends on.

Why the load order matters

Sprockets was built in an era when JavaScript expected globals: $, jQuery, _, toastr, whatever your app was using. Legacy code assumes those globals exist. If you tried to load the Sprockets bundle first, none of your modern module code has run yet, and the legacy bundle blows up trying to reference a jQuery that isn't there.

Loading the import map tags first lets you bridge from ES module land back to the global namespace before the Sprockets bundle runs. Your app/javascript/application.js looks something like this:

import jQueryModule from 'jquery'
window.jQuery = window.$ = jQueryModule

import _ from 'lodash'
window._ = _

import toastr from 'toastr'
window.toastr = toastr

By the time the deferred Sprockets bundle runs, window.jQuery is set. Any legacy plugin that starts with (function($){ ... })(jQuery) finds its argument and works.

The static import gotcha (Safari specifically)

There's one subtlety worth calling out, because it will silently bite you if you get it wrong.

Use static imports for anything the legacy code depends on. Not dynamic imports.

// This works
import jQueryModule from 'jquery'
window.jQuery = window.$ = jQueryModule
// This will silently fail in some browsers
import('jquery').then(mod => { window.jQuery = window.$ = mod.default })

Static imports resolve synchronously as part of the module's top-level execution. Dynamic imports return a promise, which schedules a microtask. The defer attribute on the Sprockets bundle doesn't wait for microtasks. It waits for DOM parsing to complete.

If the microtask hasn't flushed by the time defer runs, your legacy code executes before window.jQuery gets set, and the whole Sprockets bundle fails on the first $ reference.

Chromium browsers happen to flush microtasks before running deferred scripts. Safari is more aggressive about running deferred scripts as soon as the DOM is parsed, and it does not consistently wait for pending microtasks. Firefox falls somewhere in between depending on version. So the bug shows up as "everything works locally in Chrome, everything is broken in Safari, and the console error is a bare $ is not defined that gives you no clue what went wrong."

I learned this on the ERP client. The fix was one keyword change (import 'jquery' instead of import('jquery')), but figuring out that this was the issue took most of a day, because "works locally, breaks in Safari, error message useless" is a specific kind of pain. We now keep a comment in the actual application.js explaining why the imports are static, so future maintainers don't optimize the "clean" dynamic-import version back in without understanding what breaks.

Bridging globals for jQuery plugins specifically

The jQuery plugin world is where the coexistence pattern earns most of its keep. A typical jQuery plugin is a self-executing function wrapped around jQuery:

(function($) {
  $.fn.myPlugin = function(options) {
    // ...
  }
})(jQuery)

If jQuery isn't on window when this runs, the plugin throws immediately. Since Sprockets concatenates every included JavaScript file into one bundle, a single missing global at the top of the bundle can cascade into every subsequent plugin failing.

The reliable pattern is to set every global the legacy code depends on in application.js, in the order the legacy code expects them:

// Order matters, set jQuery first so plugins can find it
import jQueryModule from 'jquery'
window.jQuery = window.$ = jQueryModule

// Then extensions that patch jQuery
import 'jquery-ui'
import 'jquery-ujs'

// Then utility globals other legacy code depends on
import _ from 'lodash'
window._ = _

import toastr from 'toastr'
window.toastr = toastr

// Then anything that expects those globals to already be set
import 'application-legacy-bootstrap'

jquery-ui and jquery-ujs are imported for their side effects, they patch window.jQuery when they run. As long as jQuery is on the window by the time they execute, they attach themselves correctly.

For plugins that ship as UMD modules but don't publish clean ES module builds, the pattern is the same: import for side effects, let them find their globals on the window.

Where Sprockets keeps living

Once the pattern is in place, Sprockets stops being an obstacle and starts being a boring pipeline that quietly serves the legacy stuff. It's still the right home for several things.

Rails gem-provided assets. Gems like administrate, recurring_select, and older Rails engines still ship their JS through Sprockets. There's no reason to fight this. Let Sprockets serve them.

Custom app JS that isn't causing pain. If you have thirty files of application-specific JavaScript that's been stable for five years, don't rewrite it because a blog post says import maps are the future. It's already working.

CoffeeScript files. If your app has legacy CoffeeScript, Sprockets is where it lives. Import maps don't compile. This alone will keep Sprockets in the mix for many older apps for years.

Vendored JS you don't control. Third-party libraries that assume the global namespace live comfortably in Sprockets. This includes the DataTables family, most jQuery UI extensions, and anything that predates the ES module era.

Legacy CSS pipelines. Sprockets is also serving your stylesheets tree. Import maps don't handle CSS. Until you move the CSS pipeline to something like propshaft or dartsass-rails, Sprockets keeps that responsibility too.

What import maps take over

Once you've got the coexistence pattern working, import maps become the home for the modern side of your app.

Anything ES module-native. Modern libraries that ship as ES modules, no build step needed.

Stimulus and Turbo. The Hotwired stack was designed for this pattern. Adding a single Stimulus controller to a legacy page is a one-file change.

CDN-hosted dependencies you don't want to bundle. Sentry, third-party analytics, anything you'd rather pull from a jspm or jsdelivr URL and let the browser cache aggressively.

New JS you're writing today. New code goes in app/javascript/. Old code stays where it is. Over time, the balance shifts.

Coexistence as a long migration path

The interesting thing about this pattern is what happens to it over time. In practice, "side by side forever" isn't usually the destination. What usually happens is that custom JavaScript slowly migrates from Sprockets to import maps, one file at a time, as engineers touch each part for other reasons. New features go to import maps by default. Old features migrate when they're already being changed for other work. After a year or two, the Sprockets bundle has stopped growing, and eventually it stops shrinking too. It stabilizes around what it can't easily leave, which is usually just the gem-provided assets and third-party libraries that don't ship as ES modules.

That's a completely reasonable end state. The goal was never to eliminate Sprockets. It was to stop it from being an obstacle to modernization. Once it's just quietly serving the assets that belong there, it's done its job. If a future gem release ships as ES modules, you migrate that too. Otherwise you leave it alone. The pattern is coexistence as a bridge, not coexistence as a permanent architecture.

The other benefit of the slow migration is that it lets you add real JavaScript tests as you go. Every time you touch a file to migrate it from the Sprockets pipeline to import maps, that's a natural moment to add the test coverage that the original file never had. Over a year of doing this, we've added meaningful JS test coverage to code that would never have been tested if it required a dedicated testing effort. The migration budget is small enough that adding tests along the way doesn't blow it up.

When this pattern is the right answer

If any of these are true, side-by-side coexistence is probably the right move.

You inherited an app with a legacy JS pile and you're trying to add new features without rewriting the old ones.

Your team's engineering time is worth more than the aesthetic wins of a full modernization.

Your legacy JS is working. It ships. It's not the source of your pain.

You're on Rails 7 and you want import maps, Stimulus, or Turbo without introducing a bundler.

You have limited or no automated tests on the JavaScript side and can't safely do a bulk migration.

Your client won't fund a JavaScript rewrite that solves no visible business problem.

When it's the wrong answer

Side-by-side has one real cost: you're maintaining two mental models of how JavaScript gets to the browser. That's fine as an interim state. It's less fine as a permanent architecture. If any of these are true, consider going further.

Your legacy JS is the source of your pain, and rewriting it would save more time than it costs.

You're already planning a bigger overhaul that would touch the JS anyway.

Your Sprockets bundle has grown so large that page-load performance is suffering, and shipping half of it as ES modules would actually help.

You have strong JavaScript test coverage that would let you refactor safely at scale.

For most inherited-app situations, though, side by side is where the sanity is. Modernize where you're getting value. Leave the boring parts boring.

The paradigm shift for consulting on legacy Rails

The reason this pattern matters isn't a technical argument, it's a business one. Full JavaScript modernizations on inherited Rails apps are expensive. Not just in engineering hours, but in risk. Every legacy file you rewrite is a chance to introduce a bug that had been dormant for years. Every dependency you replace is a decision your future self will have to defend.

The side-by-side pattern lets you spend that budget where it's actually earning something. You can add the modern piece without paying the full modernization tax. You can defer the rewrite of the legacy bundle to a moment when there's a real reason for it. That's the shape of a healthy legacy app: not "everything is modern," but "we're modernizing the parts that need it, and leaving the rest alone."

For us specifically, this pattern has been a paradigm shift in how we approach legacy Rails engagements. Before we understood the coexistence pattern, "modernize the JavaScript" was a discrete project that had to be pitched, scoped, and funded, often with a business case we couldn't fully make. Now, "move the app to Rails 8 and let JavaScript modernize incrementally" is a background hum that happens naturally as we work on features. We don't ask the client for a JavaScript modernization budget. We just do it a little at a time, and the app gets healthier as a side effect of everything else.

For clients with a decade-old Rails app they intend to keep for another decade, that's the outcome that actually matters.


If you're maintaining a legacy Rails app and want to talk through what's worth modernizing and what to leave alone, that's the kind of work Rock Agile does. Get in touch.

Read More
Consulting Practice John Epperson Consulting Practice John Epperson

Team Augmentation Has a Leadership Problem

Team augmentation is one of the two services Rock Agile offers. I push back on it more often than I sell it. Here's why.

Team augmentation is one of the two services Rock Agile offers. We do it well. We have clients we've placed engineers with for years at a time. Some of those engagements are the best work anybody at the firm has done.

I still push back on it more often than I sell it. Because for a lot of the companies that call asking for a contractor to plug into their team, team augmentation isn't going to solve the problem they think they have. It's going to hide it.

Here's what I mean.

What team augmentation actually is

Team augmentation is when a company hires a contractor (or, in Rock Agile's case, a Rock Agile engineer) to sit inside the client's existing team. The contractor takes tickets from the same queue as the internal developers. Attends the same standups. Reports to the same project manager. From a project management perspective, they're a developer. From an accounting perspective, they're a contractor. That's the entire distinction.

The pitch is that you get a senior developer faster than you could hire one, without a full-time commitment. Which is true. And in the situations where that's the actual problem you have (genuine capacity constraint, work that will taper off in six or twelve months, a specific technical gap that's cheaper to rent than to hire), team augmentation is the right answer.

But that's not why most companies reach for it. Most companies reach for team augmentation because they can't ship what they said they'd ship, and they think adding engineers will fix it.

The one cause of headless development

There's one cause of a development team that's spinning without making progress, and it isn't developer capacity.

It's leadership.

When developers are working hard, shipping code, closing tickets, and yet the project isn't moving toward completion, that is by definition a failure of goal-setting, priority-setting, and decision-making. Those are leadership functions. If you're the CTO or engineering manager watching this happen, the fact that it's happening is diagnostic. The question isn't "how do I add more developers." The question is "what am I not doing that would let the developers I have make real progress."

Adding an augmented seat to a team that isn't making progress doesn't fix leadership. It just adds another developer to the failure. In many cases, it accelerates the failure, because now the team has more code to manage and a bigger surface area for context loss and coordination breakdown, but still no clearer goals.

And here's the darker part. When leadership doesn't want to name the failure as leadership, they name it as developer performance. The pattern is old and universal. It reliably rolls downhill. If the augmented developer isn't shipping the miracle that internal developers weren't shipping either, the augmented developer becomes the story: contractors are lazy, contractors are expensive, contractors don't understand the codebase, we should hire full-time next time. None of that is what actually happened, but it lets the underlying problem stay unexamined for another cycle.

I've watched this happen more times than I want to count. It's not a bad-actor story. Most of the leaders I've seen do it are decent people who genuinely believed they had a capacity problem. They didn't. They had a decision-making problem, and they solved for the wrong thing.

Where project-based work is different

I don't want to claim that project-based work is inherently superior. It isn't. It's harder. It requires more up-front planning. It requires the client to commit to a scope and a timeline, which is exactly the thing they were trying to avoid by hiring an augmented seat in the first place.

But that's why it works.

When you engage Rock Agile on a project instead of an augmented seat, we don't start until you've defined what "done" looks like. That forcing function is doing most of the work. If you can't tell us what done looks like, we know before we start that the engagement isn't ready. That's a favor to you. It's a lot cheaper to find out you can't articulate the goal in a scoping conversation than to find out three months in with an augmented developer.

Project-based work forces decisions. It forces goals. It forces deadlines. It forces the client to name the tradeoffs they'd prefer to leave unnamed. Every one of those is a place where team augmentation lets the client stay comfortable and where project-based work makes them uncomfortable in a productive way.

The 40/60 rule that nobody wants to hear

Here's an uncomfortable truth about what senior software development actually is: forty to sixty percent of it isn't writing code.

It's getting requirements clear enough to be actionable. It's reporting progress in a way stakeholders can react to. It's escalating problems before they compound. It's negotiating tradeoffs with people who own budgets, deadlines, and other teams whose work affects yours. It's diagnosing when a working relationship has become the actual bottleneck and doing the political work to unstick it.

Good developers can do some of these things. Rare developers can do most or all of them. This is the actual filter that separates production-grade senior engineers from people with the same title who never quite deliver.

The reason project-based work exposes this and team augmentation hides it: in a project engagement, you have to do the communication work, because you own the outcome. The 40-page system architecture document that nobody reads except the one IT guy who's floored that somebody wrote it down. The 52-page post-mortem that names what came from the previous contractors, what came from client-side negligence, and that one page you really hope nobody reads about what you did wrong. The three-page executive summary that translates the 52 pages into a decision the executive team can actually make. All of it. That's the work.

In team augmentation, the contractor doesn't own the outcome. The client's PM does. So the contractor doesn't have to do this work, and mostly can't (the political capital isn't theirs to spend). The 40/60 stays with the client, which is where it was already stuck.

The partnership illusion

There's a specific line I've heard from clients about their augmented engineers, and I've said it myself back when I was less honest about it: "You're not a contractor, you're a trusted partner."

It's a nice thing to say. It's almost never true.

Team augmentation, structurally, is an employer-employee relationship dressed up in contract-worker clothes. The augmented engineer is treated like an employee for day-to-day work assignment but doesn't have the standing of an employee for anything strategic. They're not in the room when the budget conversation happens. They're not asked about hiring or team structure. They can't push back on process decisions in the way an actual peer engineer could, because everyone knows they can be cut at contract end.

Over time, the relationship trends toward the actual power dynamic: transactional. Trust that was rhetorically claimed at the beginning of the engagement erodes without anyone noticing, because it was never structurally established. That erosion is what makes long-term team augmentation feel Sisyphean from the vendor side. You're pushing the same boulder up the same hill, and the hill is the fact that you weren't structured to be a partner even though you were told you were one.

The way to be a real partner is to own outcomes together. That's what project-based work is. That's what the contract structure of a project engagement does. It puts the vendor on the hook for a specific thing the client actually needs, in a way that creates real interdependence for the duration of the project. Trust in that relationship is earned, and it can compound. Trust in a team augmentation relationship is claimed, and it tends to erode.

When we still recommend team augmentation

To be fair to a service line we do sell: team augmentation is right in specific situations.

Genuine short-term capacity shortage. You have a defined amount of extra work over the next several months, and you can articulate what "done" means for it, but you don't want to hire full-time for a temporary need.

A specific technical gap. You need someone with a specific skill (Docker infrastructure, a Rails upgrade, mobile expertise) that your team lacks and that would take too long to develop internally.

A strategic knowledge-transfer engagement. You want a senior contractor to mentor your junior developers on a specific technology or practice, and there's a clear exit plan.

Bridge to a hire. You know you need a full-time engineer, you're actively recruiting for that role, and you need capacity in the meantime.

Notice the pattern. In each of these, the client can name the specific problem the augmentation is solving. That naming is what makes it different from team augmentation as a workaround for something else.

The honest ask

If you're considering hiring a contractor to plug into your team, ask yourself one question first: if the augmented engineer works out perfectly and ships everything they're asked to ship, will the project be done in the time you said it would be done?

If yes, you have a real capacity problem, and augmentation is the right answer.

If no, you have a leadership problem, and augmentation is going to make it worse.

That's the honest version. I've had that conversation a lot. Most of the time it saves the client from a bad engagement. Some of the time it converts into project-based work. Once in a while, the client hears it and hires someone else who's willing to be less honest with them.

That's fine. We'd rather lose that engagement than take it.


If you're staring at a team that isn't shipping and trying to decide whether to bring in outside help, it might be worth an honest conversation before you sign anything. That's the kind of scoping call Rock Agile does. Get in touch.

Read More
Consulting Practice John Epperson Consulting Practice John Epperson

Kanban vs Scrum: Which One Is Actually Right for Your Team?

Choosing between Kanban and Scrum isn't really about the work. It's about how much your team trusts each other.

The most common answer you'll hear is "it depends," which is true and also useless. So let me be more specific: it depends on how much your team trusts each other.

Both are agile processes. Both are designed to be flexible. Both are meant to give developers ownership over how the work gets done. The differences are less about ceremony and more about what the process assumes about your team's culture.

Scrum in one paragraph

Scrum organizes work into sprints, usually two-week windows. You commit to what you'll finish in the sprint, work through it, review at the end, rest, and start the next one. The rhythm is the point. Developers get autonomy inside the sprint boundary, and everyone gets a predictable cadence for checking in on whether things are on track.

Scrum works best on larger teams and on longer projects, because the sprint boundary gives you natural checkpoints to reassess priorities without renegotiating every day. It's also better when the business needs predictable delivery windows for reasons beyond engineering: sales cycles, regulatory dates, customer commitments.

Where Scrum breaks: teams under pressure to squeeze more work into each sprint often skip the rest and review beats. Sprint after sprint of stretch goals with no recovery time is how you turn a healthy Scrum team into a burned-out one. Once the trust in the sprint boundary is gone, it doesn't come back easily.

Kanban in one paragraph

Kanban is looser. Work items sit on a board with fixed columns (backlog, in progress, review, done, whatever fits your workflow). Developers pull the next item when they finish the current one. There's no sprint boundary. The team's throughput is what it is, and you measure it directly instead of guessing at it in sprint-planning meetings.

Kanban works best on smaller teams that already trust each other. Without the sprint boundary, the rhythm is set by the team's actual pace, which means it only works if everyone shows up and does the work. A high-trust team gets more done in Kanban than they would in Scrum because there's less ceremony overhead. A low-trust team drifts.

Where Kanban breaks: if trust is uneven, the fast developers end up carrying the slow ones. Nobody talks about it in standup, because there's no natural moment to. The imbalance just accumulates until somebody quits.

So which one?

Honestly, ask this: does my team trust each other to do the work without external pressure?

If yes, Kanban is going to feel lighter and get more done. If no, Scrum's rhythm will do some of the work for you until you've built the trust to move to something looser.

That's the actual question. The methodology books frame this as a choice about the work (is your work continuous or bursty, is it high-volume or high-complexity), and there's some truth to that. But in practice, the deciding factor is almost always the culture and the trust level. Choose the process that fits the team you have.

What we use at Rock Agile

We use Kanban. It fits a small, high-trust team, and it lets us respond to whatever the client needs without renegotiating a sprint plan every two weeks. If that stopped working, if we grew or the trust changed, I'd switch without ceremony. What matters is the work getting done well and the client staying happy. The process is a tool, not a religion.


If you're rethinking how your engineering team plans and delivers work, that's the kind of conversation Rock Agile likes to have. Get in touch.

Read More
AI + Engineering John Epperson AI + Engineering John Epperson

AI Doesn't Replace Discipline. It Rewards It.

I've been shipping fast with AI. Turns out the reason isn't the tool. It's the discipline I've been feeding it.

Here's an observation, held loosely: the speed of AI-assisted development is partially built on sacrificing software development discipline. The people getting the most out of AI are the ones who refused to make that trade.

I've been thinking about this because of what I've been watching in others and, on reflection, what I've been doing myself. I've been shipping fast with AI on my own projects, and my instinct on why has been "AI is a really good assistant." That's true, but incomplete. What I actually think happened is that I was maintaining strict software discipline in those codebases before I started leaning heavily on AI, and it turns out that discipline is exactly what lets AI be useful.

I was dodging a bullet without seeing it. Now I see it.

What the industry has been calling AI speed

If you've been watching the conversation, "AI is making developers dramatically faster" is the default headline. Post-hoc surveys keep reporting it. Vendor marketing amplifies it. Every LinkedIn post about AI productivity insists the multiplier is real.

Some of it is. The best AI-assisted work is dramatically faster than the equivalent solo work, I've watched it, I've done it. But "the best" is doing a lot of work in that sentence. The average AI-assisted output is not what the headlines describe. Something else is happening in the difference.

What discipline is actually made of

When engineers talk about "software discipline," they usually list the wrong things. Not docstrings on every method. Not process ceremonies. Not style guides that live in a wiki nobody reads.

The discipline that matters is the discipline that captures why. The reasoning behind the code, not just the code. That reasoning lives in a specific set of places:

  • The pull request thread where two engineers argue over an edge case and one convinces the other. The commit message is a summary; the argument is the substance.
  • The ticket comment where a product manager clarifies what "done" means for a feature and the acceptance criteria evolve.
  • The Slack conversation where someone points out a constraint from a system three teams over that nobody would have known to consider.
  • The code comment that says "this looks weird because [historical bug]" or "yes, this reads redundant, but the alternative caused problem X in production."
  • The architecture document that captures what was considered and rejected, not just what was chosen.
  • The tribal knowledge in senior engineers' heads because nobody wrote it down, but they know to ask each other before touching the auth layer.

This is the boring, often-hated stuff. Half the industry hates writing pull request descriptions. Everyone hates writing architecture documents that nobody reads. Nobody wants to type up the Slack conversation where the tricky decision got made. Most engineers view this work as tax on shipping.

But this is the work that makes future work possible. It's also, as it turns out, the work that makes AI useful.

The empirical version of the claim

I've been holding all of this as an operating theory. Two weeks ago it moved.

Markus Borg (CodeScene, Lund University) and Adam Tornhill published a peer-reviewed study on June 25, 2026 with a direct empirical version of the argument. The headline: AI coding assistants increase defect risk by 30% or more when applied to unhealthy code. Not "sometimes" and not "in edge cases." Systematically. When the study gave AI structural guidance about the codebase, refactoring success went from 5.7% of files (all code smells removed) to 52%. Without that guidance, only 24.1% of files reached what the paper calls a "human- and AI-friendly state." With it, over 90% did.

The gap between AI on an ordinary codebase and AI on a codebase with actual discipline behind it is enormous. That gap is what people are seeing but not naming when they report their AI experiences and get wildly different numbers.

Both blog and paper links at the end.

My own version of the story

I did not set out to test any of this. But looking back, I've been maintaining what I now recognize as unusually strict discipline on my personal projects for reasons that pre-date the AI wave. Old habits from the years when I was less senior and had to build compensating structure to keep pace. Real tests. Architecture notes that get updated when the architecture changes. Pull request descriptions that explain reasoning, not just contents. Comments that name the tricky decisions where they live.

When I bring AI into those codebases, it works. Not "sometimes." Not "if I write the prompt just right." Reliably. The AI can read the tests and the notes and the comments; it produces suggestions that fit the shape of what's already there; when it goes wrong, it goes wrong in visible, correctable ways because the surrounding discipline gives me signal about what a good answer looks like.

I've watched other engineers, at client codebases and open-source repos, get radically different results with the same AI tools on codebases without that discipline. Not because they're worse engineers. Because their context is starved. The AI has nothing good to pull from.

The lesson I kept missing until this month: the reason I've been high-productivity with AI is not that I'm using AI well. It's that I've been feeding it well. Discipline is what I've been feeding it.

The framing that ties it together

AI is a fantastic assistant. It is not a developer replacement.

That distinction matters. If you treat AI as a developer replacement, you're delegating the work AND the discipline. You are asking AI to produce code AND to bring the context AND to enforce the standards AND to know why the previous decisions were made. It cannot do all of that. It cannot do most of that. So you get output that looks reasonable and is subtly wrong in ways you cannot catch, because you're not bringing the context either.

If you treat AI as an assistant, the split of labor works. You bring the context. AI brings the throughput. You bring the reasoning. AI brings the transcription of that reasoning into working code. You bring the discipline. AI amplifies what discipline produced.

The people I know who are getting the most out of AI right now, without exception, work this way. They didn't start doing so because of AI. They were already doing so, and now they have more leverage per unit of discipline.

The people getting the least out of AI, or getting negative return, are treating the tool as if it can substitute for the missing discipline. It can't. It can only amplify what's already there. When the substrate is thin, amplifying it doesn't produce good code. It produces more code.

Pure vibe coding is the extreme version of this. If you're asking AI to produce features on a codebase with no tests, no comments, no reasoning captured anywhere, and no architectural thinking at the surface, you're asking for every problem that comes with that. It's not a tool failure. It's a discipline failure that the tool made faster.

What to actually do

For engineers already leaning heavily on AI:

Look at your own workflow honestly. Ask: what is the substrate the AI is pulling from when I ask it to do work? If the answer is "the codebase, and that's it," you're in the thin-substrate case. Improvements in prompt engineering will not fix that. Fix the substrate. Write the architecture notes. Add reasoning to your pull request descriptions. Comment the tricky decisions where they live in the code.

For engineering leaders deciding where to invest:

The right question is not "which AI tool should we buy?" It's "what discipline are we sacrificing for speed, and what will it cost us in the fourth quarter?" A team shipping fast on messy code with AI is accumulating debt that hits Q3 or Q4 or the next hire. A team shipping fast on clean, well-documented code with AI is compounding.

The tell is code churn. If the codebase's two-week rewrite rate is climbing as AI adoption climbs, that's not "engineers being more productive." That's the substrate getting thinner, then thinner still.

For people uncertain whether they're in the "using AI well" or "using AI badly" case:

The uncomfortable test is to take AI away for a week on a specific project and see what changes. If your code quality improves without it, your AI use was covering for missing discipline. If you slow down but your code stays about the same quality, your AI use was legitimate throughput amplification. Both cases are useful to know. Neither is a permanent judgment.

The move that follows

I've been calling this an "operating theory" throughout. That's my own hedge that I don't have controlled evidence for the causal claim. What changed this month is that the mechanism is empirically supported at the substrate level (Borg and Tornhill). The rest of my claim, that maintaining discipline is what makes AI reliably valuable, remains operating theory, but with a much stronger foundation than it had.

The prescription is small. Keep writing the pull request descriptions. Keep updating the architecture notes. Keep leaving the "here's why this looks weird" comments. Keep asking AI to explain its reasoning back to you and correcting the parts it gets wrong. These are cheap moves individually. They compound.

If you're currently getting less than you expected from AI and you're chasing better tools or better prompts, consider chasing better discipline instead. That's where I'd bet the largest gap is.

Related reading

Sources


If you're navigating AI adoption at your company and want a conversation about how the discipline side of it actually holds up under load, that's the kind of thing Rock Agile likes to talk about. Get in touch.

Read More
Ruby & Rails Craft John Epperson Ruby & Rails Craft John Epperson

The Refactor That Wasn't Broken

The most important refactors are the ones where nothing's broken yet. A walk-through of a real Ruby refactor that started with dissatisfaction, not a bug.

The most important refactors are the ones where nothing's broken yet.

I'd just landed a feature. A cross-app "What's Next" widget that surfaces things the user should do next. Incomplete onboarding. Missing documents. Ungenerated content sections. The first draft worked. Tests passed. Lint was clean. All the mount points rendered. Then I read the code again and didn't like it.

The trigger was internal dissatisfaction. Not a bug. Not a performance problem. Just a strong sense that the shape was wrong. That's a legitimate signal, and refusing to act on it because "nothing's broken" is how codebases quietly rot.

Refactors that come from dissatisfaction are the ones I trust most. Bug-driven refactors usually have a narrower scope: you're fixing the specific thing that broke. Refactors that come from a dissatisfied read of code that works are usually restructuring something that would have caused pain later, and doing it before the pain arrives. That's cheaper.

Here's how I walked the refactor from that vague sense of wrong to a shape I could defend.

Stage 0: The starting code

A 200-line service with eight item-builder methods. Two of them looked roughly like this:

def content_items
  nudges = []
  nudges << content_ungenerated_item if ungenerated_sections.any?
  nudges << content_over_cap_item    if over_cap_count.positive?
  nudges.compact
end

def content_ungenerated_item
  missing = ungenerated_sections
  labels  = missing.map { |s| I18n.t("app.content.sections.#{s}.label") }
  item(kind: :nudge,
       step: :content,
       label_key: "user.content_ungenerated",
       label_params: { count: missing.size, sections: labels.to_sentence },
       cta_url: @url_helpers.content_path,
       cta_method: :get)
end

And a couple methods later:

def document_items
  nudges = []
  nudges << document_missing_item  if missing_document_types.any?
  nudges << document_expired_item  if expired_document_count.positive?
  nudges.compact
end

def document_missing_item
  missing = missing_document_types
  labels  = missing.map { |t| I18n.t("app.documents.types.#{t}.label") }
  item(kind: :nudge,
       step: :documents,
       label_key: "user.documents_missing",
       label_params: { count: missing.size, types: labels.to_sentence },
       cta_url: @url_helpers.documents_path,
       cta_method: :get)
end

Look at those two together. They're not the same method, the domain concept differs, the I18n namespace differs, the CTA URL differs. But the shape is identical. The list-and-labels processing is the same. The item-construction is the same. The nudges << x if cond guard is the same. And the smell of copy-paste is unmistakable.

Extend that to the eight rules that were in the file, and you had eight variants of the same shape doing eight different domain jobs. That's fine when there are two of them. It's a code smell when there are eight.

And a compute method that imperatively concatenated each builder's output:

def compute
  items = []
  items.concat(onboarding_items)
  items.concat(profile_items)
  items.concat(document_items)
  items.concat(content_items)
  items.concat(export_items)
  items.sort_by { |item| WhatsNext::PRIORITY.fetch(item.kind) }
end

It worked. It passed tests. It rendered correctly on every page. But four smells:

  1. Imperative concat. items = []; items.concat(...) is "build a thing through mutation," which fights the declarative style the rest of the codebase uses. Most of this app leans on map, flat_map, and array literals with conditionals. This compute method is an island of mutation in a sea of declaration.
  2. Builder methods accepting parameters that were all derivable from self. Every keyword argument passed to item(...) was instance state. That's the data-clump smell in one of its most recognizable forms: passing the object's own state back to itself as arguments.
  3. Duplicated label processing. Two of the eight callers mapped a list through I18n.t and .to_sentence. Same shape, different namespace. This kind of near-duplication accumulates in a service class over time. Each new rule copies the pattern of the previous one, and by rule five you have four almost-identical helpers.
  4. A class doing eight things in one file. Each "rule" was a private method related to the others only by living in the same Ruby file. As long as this stayed at eight rules, the file was manageable. The moment it grew to twelve or fifteen (and it would), the file would be unreadable.

None of these were bugs. Every one of them was a design smell that would compound over time.

Stage 1: Lay out the options before touching code

The temptation when you spot a smell is to fix it inline. I resisted. I put three options on the table with code sketches for each.

Option A: Surface fix only. Convert the imperative shapes to declarative ones without changing the class structure. The compute method becomes:

def compute
  [
    *onboarding_items, *profile_items, *document_items,
    *content_items, *export_items
  ].sort_by { |item| WhatsNext::PRIORITY.fetch(item.kind) }
end

And each builder method flattens into the declarative form:

def content_items
  [
    (content_ungenerated_item if ungenerated_sections.any?),
    (content_over_cap_item    if over_cap_count.positive?)
  ].compact
end

Minimum churn. The imperative smell is gone. The other three smells remain. Deploy risk: near zero. But it doesn't address the deeper problem: this class is still doing eight things.

Option B: Rule classes with a light factory. Each builder becomes its own class with a uniform #items interface. The main service collapses to a registry:

class WhatsNext
  RULES = [
    Rules::Onboarding, Rules::Profile, Rules::Documents,
    Rules::Content, Rules::Export
    # ... and so on
  ]

  def compute
    RULES.flat_map { |klass| klass.new(@user, @url_helpers).items }
         .sort_by { |item| PRIORITY.fetch(item.kind) }
  end
end

Each rule class is tiny. Something like:

class Rules::Documents
  def initialize(user, url_helpers)
    @user, @url_helpers = user, url_helpers
  end

  def items
    [
      (missing_item if missing_types.any?),
      (expired_item if expired_count.positive?)
    ].compact
  end

  private
  # ... the domain-specific helpers ...
end

Eight tiny files, each independently testable. Pattern matches an existing precedent in this codebase (another service was already using the flat-map-a-list-of-classes pattern). Adding a ninth rule is one new file with one new line in the RULES list.

Option C: Real domain models plus presenters. Pull the "what state is this domain in" question into actual models:

class ProfileState
  def initialize(user)
    @user = user
  end

  def missing_fields
    [
      (:primary_specialty if @user.primary_specialty.blank?),
      (:bio if @user.bio.blank?),
      (:location if @user.location.blank?)
    ].compact
  end

  def complete?
    missing_fields.empty?
  end
end

Then the WhatsNext service becomes a thin presenter over the domain layer:

class WhatsNext
  def compute
    assessments = [
      ProfileState.new(@user),
      DocumentCoverage.new(@user),
      # ... and so on
    ]
    assessments.flat_map { |a| items_from(a) }
               .sort_by { |item| PRIORITY.fetch(item.kind) }
  end
end

Just describing three alternatives surfaced tradeoffs I'd have missed if I'd started refactoring right away. Option A had almost no upside beyond removing the imperative concat. Option C had significant upside, but only if we had enough evidence to justify the domain layer today. Option B was the sensible middle. Even knowing B was right, sketching A and C made me confident in it.

The principle: before refactoring, force yourself to write down 2-3 alternatives. Just describing them surfaces tradeoffs you'd otherwise miss.

Stage 2: Pick by evidence, not by aesthetic

I leaned toward B. But a sharper question came up: is C worth the cost? "Do the domain models read nicely" isn't the right frame. The right frame is: is there existing duplication or imminent reuse to justify pulling out the extra layer?

I went searching for real duplication of the ProfileState-style query across the codebase. The actual grep commands I ran looked something like:

# Find every place we're checking "is this profile field filled in?"
rg 'primary_specialty.presence' --type=ruby
rg 'primary_specialty.present?' --type=ruby
rg '@user\.(bio|location|specialty)' --type=ruby -A 2

# For the document coverage query:
rg 'user\.documents\.where\(type:' --type=ruby -B 1 -A 3

# For the content readiness query:
rg 'ungenerated_sections' --type=ruby

The output shape for the ProfileState query looked like this:

app/services/generate/resume_base.rb:23:  fallback = user.primary_specialty.presence || "(not declared)"
app/services/generate/opportunity_resume.rb:47:  fallback = user.primary_specialty.presence || "(not declared)"
app/services/generate/linkedin_section.rb:31:  fallback = user.primary_specialty.presence || "(not declared)"
app/services/score_specialty.rb:12:  return :low if user.primary_specialty.blank?
app/services/whats_next.rb:88:  missing << :primary_specialty if user.primary_specialty.blank?

Five callers. Three of them in generation services. One in a scoring module. One in the new What's Next rule.

Concrete evidence per candidate:

ProfileState. Five places already. Plus imminent reuse in three upcoming features. Verdict: extract today. High-confidence reuse and immediate duplication.

DocumentCoverage. Only one caller today. But the roadmap included a "document tailoring" feature that would need this exact query as a direct lookup rather than as an AI-inferred flag. Verdict: defer until that feature forces the second caller. Extracting it now would be architectural speculation, the correct instinct is to wait for evidence.

ContentReadiness and ExportRoster. One caller each, no clear second. Verdict: keep inline.

Along the way I made an honest correction. I'd initially claimed the four generation services were duplicating the ProfileState query. Looking again, they were doing presence || "(not declared)" as a prompt-interpolation idiom that was adjacent to but different from the completeness question. The state question is "is this field filled in?"; the interpolation idiom is "how do I display this field in prompt text?"

Look at the specific pattern used in the generation services:

# In app/services/generate/resume_base.rb:
specialty = @user.primary_specialty.presence || "(not declared)"
prompt = "You are writing a resume for a #{specialty}..."

versus what ProfileState would want to do:

# In ProfileState:
def missing_fields
  fields = []
  fields << :primary_specialty if @user.primary_specialty.blank?
  # ...
end

Both touch @user.primary_specialty. Both involve a blank/present check. But they're doing different jobs. The generation service is producing a display string. ProfileState is producing a domain classification. Conflating them would make ProfileState a god-object serving two concerns.

So: extract ProfileState for the state question. Leave the generation services alone. The four sites where presence || "(not declared)" appears will keep doing that inline, because it's a display concern in that context, not a domain concern.

The principle: "extract when a second caller appears" is a useful brake on domain modeling. It prevents speculative architecture. Domain models without callers are debt. Wait for evidence. The exception is when you can find duplication already in the code AND know the next caller is imminent. That's high-confidence reuse.

The corollary is that being honest about what "duplication" means matters. Two lines of code that look similar aren't necessarily the same concern. If you extract them together, you'll end up with a helper class that has to fork behavior based on which caller is calling it, and now you've made two separate concerns into one entangled one.

Stage 3: Naming matters more than the structure

With Option B and the ProfileState extraction agreed on, the next push came from an unexpected place: the name.

WhatsNext was a UI label leaking into the domain layer. The service's actual job was measuring the state of the workspace across multiple dimensions. Three honest framings on the table:

Framing Fit Notes
Readiness All 8 rules "Is the workspace ready in this dimension?" Covers gaps, queues, and broken artifacts uniformly.
Completeness 5 of 8 Strong for gaps, awkward for review queues and defensive items.
Audit All 8 Honest about what the service does. Compliance connotations.

Picked Readiness. It unified the meaning across all eight rules, the namespace nested cleanly, and the user-facing "What's Next" widget could stay as a thin presenter over the domain output. In code, that layering wires up like this:

# Domain layer, the reusable, business-facing thing:
module Reviewer
  RULES = [
    Reviewer::Onboarding,
    Reviewer::Profile,
    Reviewer::Documents,
    Reviewer::Content,
    Reviewer::Export
    # ...
  ]

  def self.compute(user, url_helpers)
    RULES.flat_map { |klass| klass.new(user, url_helpers).items }
         .sort_by { |item| WhatsNext::PRIORITY.fetch(item.kind) }
  end
end

# Presentation layer, the specific widget that surfaces this:
class WhatsNext
  def self.compute(user, url_helpers)
    Reviewer.compute(user, url_helpers)
  end
end

For now the presenter is a one-line delegate. But if we ever need a second presenter, an email digest of pending items, a CLI status command, a health-check endpoint that gates the workspace as "not ready", the domain layer is the reusable thing. The UI layer is what changes per surface.

The principle: when the existing name is a widget label, push back. The domain layer should be named for what it does, not how it's rendered.

Naming discussions feel like bikeshedding when you're in them. They're not. The name you pick determines what future engineers think the code is for. Choosing "Readiness" over "WhatsNext" is choosing to communicate that this code will be reused across contexts. Choosing "WhatsNext" would have communicated that it's a UI widget and would have implicitly discouraged the future presenter.

Stage 4: First pass, then a second look

First implementation: eight rule classes, each with their own #items method:

class Reviewer::Onboarding < Reviewer::Base
  def items
    return [] if @user.onboarded?
    [item(kind: :blocking, step: :onboarding,
          label_key: "user.onboarding_incomplete",
          cta_url: @url_helpers.onboarding_path, cta_method: :get)]
  end
end

Tests passed. Lint clean. Pushed.

Then I re-read the code and found more smells the first pass had left behind. Every #items was rebuilding the WhatsNext::Item structure inline. Every keyword argument to item(...) was derivable from self. Two rules were duplicating the same list-to-sentence label processing.

This was the Template Method pattern asking to be born. The Base class should own the orchestration. Subclasses should describe their properties via hook methods:

class Reviewer::Base
  LABEL_NAMESPACE = "app.workflow.whats_next".freeze

  def initialize(user, url_helpers)
    @user = user
    @url_helpers = url_helpers
  end

  def items
    applies? ? [item] : []
  end

  def item
    WhatsNext::Item.new(
      kind: kind, step: step, label: label,
      cta_url: cta_url, cta_method: cta_method
    )
  end

  private

  attr_reader :user, :url_helpers

  # Hooks, subclass overrides
  def applies?     = raise NotImplementedError
  def kind         = raise NotImplementedError
  def step         = raise NotImplementedError
  def label_key    = raise NotImplementedError
  def cta_url      = raise NotImplementedError
  def label_params = {}
  def cta_method   = :get

  # Derived helpers
  def label
    I18n.t("#{LABEL_NAMESPACE}.#{kind}.#{label_key}", **label_params)
  end

  def localized_list(keys, namespace:)
    keys.map { |k| I18n.t("#{namespace}.#{k}") }.to_sentence
  end
end

And the subclasses collapse to property declarations:

class Reviewer::Onboarding < Reviewer::Base
  private

  def applies?  = !user.onboarded?
  def kind      = :blocking
  def step      = :onboarding
  def label_key = "user.onboarding_incomplete"
  def cta_url   = url_helpers.onboarding_path
end

Five hook overrides. No construction noise. No parameter passing. No return [] unless ... boilerplate. Each method does one thing. The grain is right.

For rules with computed label parameters, the same pattern holds:

class Reviewer::Profile < Reviewer::Base
  private

  def applies?     = user.onboarded? && !state.complete?
  def kind         = :nudge
  def step         = :profile
  def label_key    = "user.profile_incomplete"
  def cta_url      = url_helpers.profile_path

  def label_params
    { fields: localized_list(state.missing_fields,
                             namespace: "#{LABEL_NAMESPACE}.profile_fields") }
  end

  def state
    @state ||= ::ProfileState.new(user)
  end
end

And for a rule that has multiple ways to trigger:

class Reviewer::Documents < Reviewer::Base
  private

  def applies?  = missing_types.any? || expired_count.positive?
  def kind      = :nudge
  def step      = :documents
  def label_key = missing_types.any? ? "user.documents_missing" : "user.documents_expired"
  def cta_url   = url_helpers.documents_path

  def label_params
    return { count: missing_types.size } if missing_types.any?
    { count: expired_count }
  end

  def missing_types
    @missing_types ||= DocumentCoverage.new(user).missing_types
  end

  def expired_count
    @expired_count ||= user.documents.expired.count
  end
end

The label_key and label_params are conditional because the same rule handles two related states. That's fine, the rule is small enough to hold both conditions without the class becoming confused. If the branching got more complex, that would be a signal to split into two rules.

Each method does one thing. The grain is right. And adding a ninth rule is one file with five hook overrides.

Ruby techniques worth calling out

A few Ruby-specific details make this pattern read especially well:

Endless method definitions (def applies? = !user.onboarded?). Ruby 3.0+. They make hook overrides read as data rather than as code. Compare:

# Regular form, six lines per property
def applies?
  !user.onboarded?
end

def kind
  :blocking
end

To:

# Endless form, two lines
def applies?  = !user.onboarded?
def kind      = :blocking

Same behavior. The endless form makes the "declaration" nature of the code visible. When every subclass's overrides are single-expression declarations, the endless form removes the def...end noise and lets the properties themselves become the visual content of the file. If you're glancing at a rule class, you should be able to see its full property definition in a screen. Regular def...end blocks make that hard for a class with five hooks; endless methods make it easy.

attr_reader :user, :url_helpers in the base. Small but meaningful. Compare:

# Without attr_reader, subclasses use instance variables
def applies?  = !@user.onboarded?
def cta_url   = @url_helpers.onboarding_path

To:

# With attr_reader on the base, subclasses use methods
def applies?  = !user.onboarded?
def cta_url   = url_helpers.onboarding_path

The second reads as a property declaration ("I need user to check whether they're onboarded"). The first reads as instance-variable access ("I need @user, but @user where? Set where? Any subclass can accidentally reassign this."). The attr_reader also enforces read-only access to the constructor arguments, which prevents a whole class of bugs where a subclass would clobber @user mid-execution.

The declarative array-with-conditionals form. Compare:

# Imperative, build through mutation
def items
  arr = []
  arr << missing_item if missing_types.any?
  arr << expired_item if expired_count.positive?
  arr
end

To:

# Declarative, describe what the array can contain
def items
  [
    (missing_item if missing_types.any?),
    (expired_item if expired_count.positive?)
  ].compact
end

Same output, different reading order. The first describes an algorithm ("start empty, conditionally push, return"). The second describes a specification ("here are the two items this rule can produce"). In a service class whose job is to describe eligibility, the second form reads as the answer to the question. The first reads as the process by which the answer gets computed. When you're describing rules, the specification form is almost always right.

flat_map and sort_by as a one-liner pipeline vs. imperative concat-and-sort. Compare:

# Imperative
def compute
  items = []
  RULES.each do |klass|
    items.concat(klass.new(user, url_helpers).items)
  end
  items.sort_by { |i| PRIORITY.fetch(i.kind) }
end

To:

# Declarative pipeline
def compute
  RULES.flat_map { |klass| klass.new(user, url_helpers).items }
       .sort_by { |i| PRIORITY.fetch(i.kind) }
end

The chain is two operations that compose. The imperative version is five lines that don't compose. In Ruby, whenever you can express a transformation as a pipeline, it reads better than as a loop with mutation.

Testing the pattern

Each rule class is independently testable, which is one of the biggest reasons for extracting them. Before the refactor, testing the WhatsNext service required constructing a User in various complete/incomplete states and asserting on the aggregate output. After the refactor, testing a rule is a two-line stub:

RSpec.describe Reviewer::Onboarding do
  let(:url_helpers) { double(onboarding_path: "/onboarding") }

  it "returns a blocking item when user is not onboarded" do
    user = build_stubbed(:user, onboarded: false)
    rule = described_class.new(user, url_helpers)

    expect(rule.items.size).to eq(1)
    expect(rule.items.first.kind).to eq(:blocking)
    expect(rule.items.first.step).to eq(:onboarding)
  end

  it "returns no items when user is already onboarded" do
    user = build_stubbed(:user, onboarded: true)
    rule = described_class.new(user, url_helpers)

    expect(rule.items).to be_empty
  end
end

The rule's tests describe the rule's contract, "here's when it fires, here's what it produces." No user state gets tangled up with other rules. The WhatsNext service, meanwhile, gets a much shorter test suite that just checks the composition, "when three rules fire, we get three items in the right order." Each layer's tests describe that layer's job.

Before the refactor, the service class had eighteen tests, all of which had to construct a User in specific combined states to exercise the various branches. After the refactor, the same behavioral coverage lives in twenty-five tests split across the rule classes and the WhatsNext composition. Each test is smaller. Each is easier to write. And when a rule changes, only that rule's tests need updating.

The dependency tree that made this possible

One thing I glossed over: this refactor was cheap because we already had a complete dependency map for the What's Next pattern. Every rule, every downstream consumer, every follow-up item that would need to change if a rule's shape changed, all mapped out and documented. Before I typed the first character of the refactor, I knew exactly what depended on what.

That's not typical. In most codebases, identifying the dependency tree for a feature like this is something you do late in the game, usually after a refactor has started to go sideways, and you're scrambling to figure out what else you're breaking. We do it upstream. The dependency map is a first-class artifact, not something that lives in developers' heads and gets discovered as needed.

The immediate value of a complete dependency map is that it makes evidence-based refactoring easy. When I said "grep the codebase for this pattern," the grep was the confirmation, not the discovery. The dependency map already told me ProfileState had five callers in specific files. The grep just verified the exact line numbers. That inverts the usual investigation loop: instead of guessing what might depend on something and then searching to verify, you already know, and searching is bookkeeping.

The deeper value is that a complete dependency map makes it easy to identify follow-up items, the places where a related architecture should look the same but doesn't yet. Once we agreed to extract ProfileState as a domain model, the dependency map told us exactly which other reviewer rules would benefit from the same treatment when their evidence arrived. Not "someday when we notice." Right now, in the docs, with the specific criterion (a second caller) that would trigger each extraction. When DocumentCoverage becomes justified by the second caller landing, the extraction won't be a fresh design exercise. It will be running the pattern we already documented.

This is a discipline most codebases skip. Their dependency maps live in the developers' heads, discovered as needed. The result: refactors are expensive because you don't know what you're going to break, so you refactor conservatively, small changes, minimum churn, leave the rest of the codebase alone. Which means the rest of the codebase stays in the pre-refactor shape indefinitely. It never gets the same treatment. The codebase drifts into inconsistency because each part gets refactored on a different schedule based on when it starts causing pain.

Having the dependency map up front inverts this. We can refactor systematically because we know the whole picture. When we improve the shape of one rule, we can immediately improve the shape of related rules that share the same architecture pattern. The refactor doesn't leak into surprise breakage because there aren't surprises. And when we sweep the codebase months later to check that all rules follow the same pattern, we're checking against a documented spec, not against a shifting collective memory.

For teams that haven't done this: it's worth doing. The upfront cost of mapping the dependency tree is real, but it pays off the first time you need to refactor anything non-trivial in that part of the codebase. And it pays off continuously, every subsequent change is done with the full picture in view rather than with the local view of just the code you're touching.

Principles distilled

Six principles I'd carry from this refactor to the next one.

The most important refactors are the ones where nothing's broken. Dissatisfaction is a legitimate input. The smell instinct is faster than the bug instinct, and codebases that only get refactored when something breaks accumulate quiet debt.

Write down 2-3 alternatives before changing code. Even when you know option B is right, sketching A and C makes you confident in B. Sometimes the sketching surfaces a better hybrid you hadn't seen.

"Extract when a second caller appears" is a useful brake on domain modeling. Domain models without callers are debt. Wait for evidence. The exception is when you can find duplication already in the code and know the next caller is imminent. That's high-confidence reuse.

Push back on names that leak presentation into the domain layer. WhatsNext was the widget label. The service was doing Readiness assessment. The rename made the layering obvious.

Be honest mid-refactor about your earlier claims. I claimed five callers were duplicating the ProfileState query. On re-read, four were doing an adjacent-but-different pattern. Saying that out loud is the difference between honest design and motivated reasoning.

Hook methods over parameters when the data is on self. Parameter lists in Ruby OO are often a sign the data should be moved onto the receiver. Subclasses that declare def kind = :blocking read like property declarations. That's the grain the language wants.

The final shape

Eight rule classes, each 12-29 lines. A Reviewer::Base at 72 lines doing all the orchestration. A Reviewer registry at ~20 lines. A ProfileState domain model at ~30 lines. Twenty-five test examples covering the rules and the new model.

Adding the ninth rule is one file with five hook overrides.

None of that was necessary. Nothing was broken. The dissatisfaction was the whole trigger. That's the point.


If you're navigating a Rails codebase where the design has quietly drifted and you're not sure where to start, that's exactly the kind of work Rock Agile does. Get in touch.

Read More
AI + Engineering John Epperson AI + Engineering John Epperson

The U-Curve of AI Amplification

AI amplifies juniors and seniors most; mid-level engineers gain the least. A counterintuitive shape with real implications for how you hire.

A working theory about who AI actually helps at work. Byline: John Epperson. Audience: CTOs and engineering leaders making org-level AI adoption decisions.


Here's an observation I keep coming back to, held as a strong opinion loosely: the value AI provides to a developer at work isn't distributed evenly across experience levels. It's not proportional to skill. It's shaped more like a U.

Junior developers get enormous value. Mid-level developers get modest value. Senior developers get enormous value again. If you plot the value against career stage, the curve dips in the middle.

That's counterintuitive. Most people assume value scales linearly with skill, or that AI democratizes expertise so it flattens the curve. I don't think either is right, and this piece is my attempt to lay out why. The stakes are practical: if the shape of the curve is real, it changes how you should hire, how you should structure teams, and where you should invest training budget.

One caveat up front, and I want to keep saying it as the piece goes on: everything below is grounded in observation, not empirical research. I'm speaking with the authority of someone who has watched a lot of engineers work with AI over the last two years across multiple codebases, not with the authority of a controlled study. When I say "juniors get more value" or "mid-level engineers plateau," I'm generalizing from cases I've watched play out. More evidence one way or another is needed. Treat this as an operating theory to test, not a claim to plan around.

What juniors get

For someone new to the field or new to a technology, AI closes knowledge gaps that used to be closed by mentorship, by Stack Overflow, by six months of pattern-matching on someone else's code. That closure used to be expensive. It's now cheap.

A junior developer with a good AI setup can ask "what's the idiomatic Rails way to do this" and get a decent answer in seconds. They can ask "why is this test failing" and get an explanation that would have taken hours to find. They can ask "explain this codebase to me" and get a map. The AI is functionally a patient guide, available 24/7, without the social cost of asking someone senior "the same question again."

The failure mode for juniors is uncritical acceptance. The AI is often wrong, and the junior lacks the intuition to catch the wrong parts. Codebases can drift into plausible-but-wrong architectures if a junior takes AI output as gospel. This is a real risk, and it's why "let juniors use AI unsupervised" is a mistake in most organizations.

But the ceiling for value at this tier is high. If the failure mode gets managed (through code review, through paired seniors, through cultural norms that treat AI as a first-draft tool), juniors go from "learning slowly" to "learning at a pace that would have taken years." That's a real transformation, and it's what most of the "AI democratizes coding" claims are pointing at when they're not overselling.

What mid-level engineers get

Here's the part that surprises people: mid-level engineers get the least from AI.

Not zero. But less than either end of the curve.

The reason is that mid-level engineers are in an awkward spot. They have enough experience that they don't need the "here's how Rails works" guidance a junior gets. But they don't yet have the deep intuition that lets a senior spot a smell and articulate it in a way the AI can amplify. Mid-level engineers use AI mostly as a rubber duck, a colleague to talk problems through. That's useful. But the leverage is capped at their own judgment level.

Concretely: a mid-level engineer will accept the AI's first-pass output more often than a senior will. Not because they're lazy, but because they don't yet have the calibrated dissatisfaction sense that says "this passes tests but the shape is still wrong." They ship the passable version. The output looks like solo work, just faster.

I watched this play out recently on a project where a mid-level engineer was working on a complex feature. They'd hand a problem to the AI, get a solution, apply it, test it, and move on. On paper, they were productive. In practice, they were spending a lot of cycles going in circles. The AI would produce a solution that had a subtle design flaw. The engineer would apply it. A day later, the flaw would surface as an edge-case bug. They'd hand the bug to the AI, get a patch, apply the patch, move on. A day later, the patch would create a different subtle problem. The engineer was coding themselves into corners they didn't realize they were building, because the AI wasn't telling them "the shape of what you're doing is off." It was just answering the immediate question they asked.

A senior looking at the same starting problem would have caught the design flaw before the first solution shipped. Or would have pushed back on the AI's first pass and asked for alternatives. The mid-level engineer didn't know to do either, and the AI didn't do it for them. So the same mid-level engineer who used to write reasonable code slowly was now writing reasonable-looking code faster, and the underlying design was drifting further from correct every day.

Faster is not nothing. But it's less than the transformation the tier above and the tier below get. And in the pathological case it's actually negative. The engineer is producing more code, more of it is subtly wrong, and the total time-to-a-correct-solution is longer than it would have been without the AI.

There's also a separate risk at this tier that's worth naming. Mid-level engineers who use AI heavily can plateau in a way they wouldn't have without it. The parts of their career where they'd naturally develop deeper intuition, the frustrating, slow, deeply-thought moments, are the moments AI is most tempting to short-circuit. Sit with the frustrating problem, work through it, feel the pattern become internalized. Or hand it to AI and get an answer. The second path is faster today. It also skips the exact experience that would have moved the engineer from mid-level to senior. That's speculative. I don't have data. But it's the failure mode I'd watch for.

What seniors get

The senior mechanism is different in kind, not just in degree.

A senior developer working on a problem has trained pattern-matching that says "something is off here" often before they can say why. That pre-verbal signal is the thing that separates senior work from mid work. The signal is right most of the time. Articulating what it's actually detecting is what takes the time.

AI compresses that articulation cost. The senior says to Claude, "this class seems to be merging arrays upward, which feels wrong." The act of naming it forces the intuition into words. Claude then proposes options, sketches alternatives, scans the codebase for supporting evidence. The senior evaluates, picks, iterates. What used to take hours (the sketching, the evidence-gathering, the tradeoff analysis) collapses into minutes.

The intuition stays in the senior's head. That's what the AI can't provide. But the throughput of investigating and acting on the intuition goes up dramatically. A senior with AI can do in a day what would have taken a week solo, at the same or higher quality.

Even seniors have failure modes with AI, though, and I want to be honest about mine. I catch large code shapes going wrong pretty reliably, I'll notice when a service class is turning into a god-object, when a controller is doing more than it should, when a domain layer is missing. That's the calibrated intuition doing its job.

But I'm less reliable at catching small smells. A method that's slightly too long. A variable name that's carrying too much semantic weight. A test that's curve-fitted to the current implementation rather than testing the behavior. Those come at me from the AI's output constantly, and my pattern is often the same: I read the code, notice the smell, feel a small friction, immediately generate a justification for why it's fine in this specific case, and move on. Some of those justifications are legitimate. The pattern really is fine in context. But some are motivated reasoning. I'm choosing not to fix the smell because fixing it would slow me down.

The seniors who get the most out of AI are the ones who've learned to notice this pattern in themselves and push through it. Fix the small smell. Refactor the slightly-too-long method. Don't accept the curve-fitted test. The productivity win from AI isn't "more code faster." It's "more code faster AND at higher quality than solo," and the higher-quality half only holds if the human keeps refusing the small smells.

Even careful code review lets things through. It always has, even before AI. What's changed with AI is the volume. Twenty PRs a week, each with a plausible-but-slightly-off pattern, adds up faster than twenty PRs a week each written from scratch by a human who was also going to make some mistakes. The percentage of subtle problems per PR might not be higher with AI; the volume of PRs is higher, so the absolute number of subtle problems reaching the codebase is higher. And once a subtle problem is in, friction increases, churn increases, and bad modeling and poor architecture compound. The surface area of things that could go wrong increases, and the future busy work to fix them accumulates.

Being senior doesn't eliminate this. It reduces it. And the reduction, over months and years, is the whole ballgame. If a senior with AI reduces the rate of new subtle problems by even a moderate percentage compared to a mid-level engineer with AI, the difference in accumulated codebase health after a year is enormous. That translates directly into productivity, the team not fighting subtle problems from six months ago is the team shipping new features today.

The GitClear signal

There's some empirical data pointing in the same direction. GitClear ran a study analyzing about 211 million lines of structured code change data from 2020 through 2024, comparing the maintainability patterns of AI-generated code against pre-AI baselines. The findings are directionally what the U-Curve theory would predict.

The proportion of new code that got revised within two weeks of its initial commit grew from 3.1% in 2020 to 5.7% in 2024. In AI-heavy projects specifically, GitClear observed roughly a 39% higher churn rate. More code was being written; more of it needed immediate revision.

Duplicated code blocks rose about eightfold in 2024 compared to previous years. Copy-pasted lines climbed from 8.3% to 12.3% of new code, while the share of code that got refactored (rather than newly written or copied) collapsed from about 25% down to under 10%. The AI's habit of regenerating similar shapes rather than reusing existing ones is measurable in the commit data.

GitClear also found that AI-authored pull requests carried roughly 1.7x more issues per PR (about 10.8 versus 6.4) than non-AI PRs in the same codebases, and that estimated technical debt increased somewhere in the 30-41% range post-adoption on the projects they analyzed.

The full report ("AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones") is publicly available if you want the methodology and the raw numbers.

These aren't "AI is bad" numbers. They're "AI amplifies whatever pattern the person driving already has" numbers. A senior engineer using AI produces measurably different output than a mid-level engineer using the same AI, and the compounding difference over months and years is what the U-Curve theory tries to capture. If AI-heavy projects see 39% higher churn on average, the projects driven by seniors are pulling that average down, and the projects driven by mid-level engineers without senior oversight are pushing it up.

For contrast, Stack Overflow's 2025 developer survey found 51.6% of respondents reported positive productivity impact from AI tools. That's a self-report number, though. The gap between "developers feel more productive" and "the codebase shows more churn and duplication" is exactly the tension the U-Curve tries to explain. Both can be true simultaneously if the productivity gain is uneven, concentrated at the seniors who are producing better code faster, offset by mid-level engineers producing more code that needs more revision.

That churn shows up in codebases too, not just in study aggregates. When a mid-level engineer uses AI heavily and doesn't have the senior instinct to catch the wrong shape early, their commit history looks busier than a senior's using the same AI. More small commits fixing small problems. More reverts. More "actually, do it this way" corrections. Each individual correction is small; collectively they represent time the engineer spent because the first pass wasn't quite right.

Common flaws AI has when coding

The mechanism of the mid-level failure mode is worth being concrete about. When I watch what AI defaults to, the patterns show up over and over:

Early assumption of correctness. The AI states its solution with confidence. Even when it's wrong. Even when the code it produces has an obvious bug. This confidence is contagious. An engineer without strong pattern-matching absorbs it and stops questioning the solution before verifying it.

Curve-fitted tests. Ask AI to write tests for the code it just wrote, and it often produces tests that match the specific behavior of that implementation rather than tests that describe the intended behavior. The tests pass. They also don't catch the case where the implementation is wrong, because the tests were written to match the implementation, not the specification.

Bloating third-party libraries. AI reaches for well-known libraries by default. Sometimes that's right. Sometimes it introduces a large dependency to do something the standard library could handle in ten lines. Left unchecked, the dependency count grows every session.

Poor readability by senior standards. The AI's default code is often mid-level-quality: it works, it's correct, it's neither offensively verbose nor cleverly compressed. Which means it lacks the specific stylistic sharpness that experienced engineers reach for. Function names that are slightly too generic. Variable names that mix abstraction levels. Comments that explain what rather than why.

Duplication. The AI generates similar shapes in different places without noticing that they're similar. Two rules that follow the same three-step structure get written as two independent methods with three steps each, instead of as a base class with the shared structure and two variants.

Each of these is small. Individually, each is fixable in a code review. Collectively, they add up to a codebase that requires more maintenance than the same codebase without them.

Keeping them in check is a chore. It requires the reviewer or the author to actively look for each pattern and push back. And this is where the senior/mid-level split matters again: the senior's calibrated intuition catches these more reliably than the mid-level engineer's. Not because the mid-level engineer is careless. Because the senior has seen the accumulated cost of each of these smells in past codebases and pattern-matches against them fast, whereas the mid-level engineer is still learning what the accumulated cost looks like.

Why the shape is a U

The reason juniors and seniors both gain is that both have a specific mechanism the AI amplifies well.

For juniors: knowledge gaps. AI fills them.

For seniors: articulation cost on intuition. AI compresses it. And the senior's judgment on what to accept and what to fix is what keeps the AI's amplified output from becoming amplified debt.

The mid-level tier is between mechanisms. They have enough knowledge that knowledge-filling matters less. They don't yet have the deep intuition that articulation-compression amplifies. Neither mechanism dominates, so the value they extract sits in the middle: real, but modest, and potentially negative in the specific case where the mid-level engineer amplifies subtle wrongness faster than they'd have produced it solo.

What this implies for organizations

Take this section with the same "operating theory" hedge as the rest. I'm reasoning from my own experience and from the patterns I've watched at client organizations. The prescriptions below are what I'd bet on, not what I've proven.

Don't hire mid-level engineers as your default. The classic hiring pipeline optimizes for mid-level because they're "safe." That was correct advice when the value each tier extracted from tooling was roughly proportional. It's less correct now. If AI amplifies juniors and seniors more than mid-level, then a team weighted heavily toward mid-level is systematically underinvesting in the tiers where AI is doing the most work.

Invest in senior mentorship for juniors. Juniors with AI get transformed. Juniors with AI and a senior watching their work get a career on rails. The risk that AI produces plausible-but-wrong code lands hardest on juniors, and the intervention that reduces the risk is a real human catching it. This is a place where mentorship pays off dramatically, and where losing it costs a lot.

Take seniors' AI concerns seriously, but push back on refusal to use it. The pattern I've seen is that seniors who won't use AI often frame it as principle ("I want to think for myself") when the real objection is discomfort with delegating any part of the work. That's fixable, but only if you name it. A senior who never uses AI in 2026 is systematically slower than one who does, and the gap grows every quarter.

Watch for mid-level plateaus. If your mid-level engineers are using AI heavily but not showing the growth trajectory you'd expect, ask whether they're being asked to do the kinds of problems that build senior intuition. Those problems are frustrating, slow, and easy to short-circuit with AI. They're also the ones that build the skill. Consider deliberately assigning some problems to mid-level engineers with an explicit "solve this without AI first, then compare your solution to what AI would give you" instruction. The comparison teaches them what to notice, which is what they need to develop the senior instinct.

Track code churn as a signal. If your codebase's churn rate is climbing as AI adoption climbs, that's a signal that the AI amplification is happening in the wrong direction, more code being written faster, but more of it being revised. If churn is stable or dropping while shipped features climb, that's the healthy pattern. This is a leading indicator you can watch monthly.

What I'm still unsure about

The domain generalization. This claim comes out of my Ruby/Rails and multi-language client work. I'm not confident it holds identically for greenfield mobile development, or for infrastructure work, or for research-adjacent domains. The shape might be different when the "intuition" the senior has is different in kind.

The mid-level definition. "Mid-level" is a fuzzy category. Some mid-level engineers have deep specialization that gives them senior-tier intuition in one subfield. They probably get senior-tier amplification in that subfield. The tier boundaries are messier than the piece implies.

The long-term cognitive question. Does heavy AI use weaken a developer's independent thinking over time? I don't know. I don't have enough evidence either way. I'm suspicious of both the "AI will atrophy our brains" and the "AI is pure upside" answers. If I had to guess, I'd say it depends heavily on how it's used, and the same tool can build or erode skill depending on the discipline the user brings. But that's a guess.

The whole shape. I keep saying "operating theory" because I mean it. This is a hypothesis to test, drawn from observation, not from data. Somebody with actual metrics, churn rates by developer, PR-review-comment counts, feature-completion times, could probably test whether the U-curve holds in their organization. I'd love to see that data. I don't have it.

If your organization is figuring out how to think about AI investment across your team, I'd rather you sit with the U-Curve as an operating hypothesis to test than as a claim to plan around. But I'd bet on the shape.

Update, 2026-07-06

An edit added after publication. On July 5th I ran across empirical research that directly bears on this post's mechanism claim, and it's substantial enough to belong in the piece rather than left for a follow-up.

Markus Borg (CodeScene, Lund University) and Adam Tornhill published a peer-reviewed study on June 25, 2026 that measured AI coding assistant performance across codebases of different health levels. The headline finding: AI coding assistants increase defect risk by 30% or more when applied to unhealthy code, and the real-world risk in legacy systems is likely higher. When AI was given structural guidance about codebase health, refactoring success went from 5.7% of files (all code smells removed) to 52%. Without that guidance, only 24.1% of files reached what the paper calls a "human- and AI-friendly state." With it, over 90% did.

That is the direct mechanism I was speculating about. The U-Curve implies that seniors and juniors extract more value from AI than mid-level engineers, and I argued the reason is that both have specific conditions AI amplifies well while the mid-level tier sits between mechanisms. Borg and Tornhill's data supports a specific version of that: the value AI provides depends heavily on the health of the code and the discipline around it. Seniors keep their code healthy and enforce that discipline; juniors work in scoped, greenfield-adjacent contexts where health matters less; mid-level engineers often end up inheriting the largest complexity in the messiest codebases, which is exactly the condition where the paper says AI introduces the most risk.

This doesn't test the U-Curve directly. Nobody has run "AI amplification value by career tier, controlled." But it does anchor the mechanism to real numbers, and it moves the U-Curve from "operating theory" toward "operating theory with an identifiable causal thread." I'd still bet on the shape.


If you're navigating AI adoption at your company and want a second opinion from someone who's been using it heavily in production for the last two years, that's the kind of conversation Rock Agile likes to have. Get in touch.

Read More
Consulting Practice John Epperson Consulting Practice John Epperson

Inheriting a Dev Project: How to Not Blow It in the First Two Weeks

The shape of the work when you inherit somebody else's codebase, how to earn the right to change it before you change anything.

At Rock Agile, we do this a lot. Someone hands us a codebase they didn't write. The original developer is gone. The documentation is stale. The tests may or may not run. The business owners want progress soon, ideally last week. This is one of the specific things we're hired for, and after years of doing it, I've learned there's a shape to the work that keeps you out of trouble.

The shape has one non-negotiable at the front: you have to check yourself before you touch anything.

Start with your mindset, not the code

Developers are naturally judgmental. We look at somebody else's code and our first instinct is to see what's wrong with it. Sometimes the code deserves the judgment. More often, we're just missing context the original developer had, and our judgment is going to blind us to what the code is actually doing.

There's a Henry Ford story I keep coming back to. Ford would take job candidates to dinner. If they seasoned their food before tasting it, he wouldn't hire them. He wasn't afraid of change, he changed things constantly, but he wanted leaders who would understand what was on the plate before they started rearranging it.

The principle transfers to software. Understand what the code is doing, and why, before you decide it's wrong. Prejudgment on an inherited codebase is the single most common mistake I see, and it usually looks like premature refactoring or optimization on parts of the code we haven't earned the right to change yet.

Curiosity works better than judgment. Come in expecting to be surprised by what you find.

Understand what the software is supposed to do

Before you dig into how it works, understand what it's for.

The best case is a product owner who can walk you through the business logic. If you have them, use them. Ask why the software exists, what problem it was originally built to solve, and how that problem has evolved. Understand the constraints that shaped the original architecture, even the ones that no longer apply.

If the product owner isn't available, the next best source is any developer who worked on it before. Ask what was tried, what worked, what didn't, and what problems keep coming back. Repeated bugs and repeated attempts to fix the same thing usually point to something structural that the previous team knew about but couldn't get to.

If nobody is available, the code has to teach you. That's slower, but not impossible.

Get a working environment before anything else

Before you form opinions about the code, get it running. Bootstrap it. Run the tests. Click through the features. Take notes on what works and what doesn't. Test in multiple environments if you can. Differences between platforms are often where the most interesting bugs live.

The specific move I make early on a Rails inherited project: find the biggest files. That's usually where the core domain lives, or where things went wrong. Read the dependency graph next. The Gemfile tells you what problems the original team decided to outsource, and outsourcing decisions age faster than internal code.

Set a timer on rabbit holes. Early in my career I gave myself five minutes on any tangent while getting a new project running. That habit still runs in the background now. I'm faster at recognizing which threads are worth pulling and which aren't. It's a discipline worth building.

Work with the existing decisions before you change them

Now you know what the software is for and how it runs. Only now should you start considering changes.

Software development is trial and error, and it's a lot more of an art than most engineers admit. My rule for inherited projects is to prioritize working with the previous team's choices before replacing them. Sometimes the choices are wrong (plenty of them will be) but you'll spot the actually-wrong ones faster if you're not fighting the whole codebase at once. And sometimes what looked wrong at first glance turns out to have been the right call for constraints you didn't know about.

When you do start making changes, keep them small and reversible. The goal in the first weeks is to build enough understanding of the codebase that you can commit to bigger changes with confidence.

Take notes obsessively

I cannot overstate this. Write down what you tried, what worked, what didn't, and what confused you. This is partly for your future self: a month from now, when something breaks, you'll want to know what you were thinking today. But it's also for whoever comes after you. Somebody will inherit this project from you eventually. The notes you take now become the documentation they'll wish existed.

Share your findings

Software isn't a solo sport. Share what you've learned with the client, with your team, with anyone else who touches the system. Other people will see things you missed. That's what makes the process work.

The through-line

The pattern is: discovery before decisions. Every step above is really about earning the right to change the code. The developers who blow up inherited projects skip the discovery. They come in confident, refactor too much too soon, and then can't tell whether their changes broke something existing or exposed a bug that was always there. The developers who inherit projects well are the ones who move slower at the start and faster later, because by the time they're touching things aggressively, they know what they're touching.


If you're looking at a legacy codebase you inherited, or one your team inherited, that's the kind of work we do. We come alongside your team, do the discovery work, and either help you build on what's there or rescue what needs rescuing. If that sounds relevant, get in touch.

Read More