Legacy code refactoring on a live .NET 8 API, with zero contract changes
How AI-native development cleared a blocking SonarQube quality gate and 26 log injection vulnerabilities in a production .NET 8 API, with zero contract changes across all 275 endpoints.
Case studyOn this page16 sections
- Why the SonarQube quality gate was blocking every release
- Our approach: AI-native development with multi-agent workflows
- Scoping the technical debt remediation with analyzer data
- Characterization tests and safety nets before any refactor
- Fixing in risk-ordered, reviewable pull requests
- Fixing 26 log injection vulnerabilities with one sanitizer
- Choosing the behavior-preserving fix, every time
- What the safety nets caught
- Results
- How to reduce technical debt without breaking production
- FAQ
- How do you reduce technical debt without breaking production?
- What are characterization tests?
- How do you fix a failing SonarQube quality gate?
- How do you prevent log injection in .NET?
- What is AI-native development?
We used AI-native development to refactor legacy code in a production .NET 8 API that bills real customers. The work cleared a SonarQube quality gate that had blocked every release, fixed 26 log injection vulnerabilities, and added more than 160 safety-net tests. None of the API's 275 endpoints changed.
Why the SonarQube quality gate was blocking every release
The platform's core API is a mature ASP.NET Core service. It handles authentication, company and user management, time tracking, reporting, payroll and Stripe billing for paying customers. Every change to the release branch was blocked by a failing SonarQube Cloud quality gate.
Reliability rated C, with 12 bugs, mostly missing request cancellation handling in middleware, plus an Entity Framework default value type mismatch.
Security rated B, with 26 log injection vulnerabilities, where user-supplied values could forge log entries.
A strict zero new issues rule counted 376 open issues against the release. The analyzer's new code window had quietly grown to cover about two years and roughly 65,000 lines of code.
The dashboard showed nearly 1,500 issues, which made the problem feel overwhelming. The constraint was also clear: this is a live system billing real customers, so nothing could change for API consumers, the database, or production behavior.
Our approach: AI-native development with multi-agent workflows
We ran this engagement as an AI-native development project, with Claude Opus 5.5 as the engineering engine and our .NET engineers setting direction and approving every change.
One lead agent, several specialists. A lead Claude Opus 5.5 agent pulled the analyzer data, built the remediation plan and wrote the safety-net tests. It then split the work into independent streams and handed each to a specialist agent: API layer, core business services, infrastructure and shared libraries, and later moderate complexity, high complexity and parameter count refactors.
Parallel work without collisions. Each specialist agent worked in its own isolated git worktree, with its own build and test run. Three streams could run at the same time without stepping on each other.
Guardrails built into every task. Each agent received explicit rules. Runtime behavior must not change. Tests must be written and passing against the original code before any refactor. The new code must not introduce new analyzer issues. Anything that can't be proven safe must be skipped and reported, not guessed at.
Trust, but verify. AI code refactoring is only as good as its verification, so the lead agent never took a specialist's report at face value. It re-read the riskiest diffs, re-ran the full build and test suite itself, and checked each pull request against the quality gate. When an analysis flagged something the agents had introduced, it fixed it within the same session.
Humans stay in the loop. Engineers approved the plan, owned the strategic decisions, such as fixing everything in code rather than suppressing warnings, and reviewed and merged every pull request.
This let a small team take on a codebase-wide remediation that would normally take weeks. Nothing was shortcut: every change was covered by tests and verified automatically.
Scoping the technical debt remediation with analyzer data
We pulled the analyzer's data directly from its API. That separated the whole-project debt, about 1,500 issues, from the 376 issues actually blocking the gate, grouped by rule, file and gate condition. The problem went from rewrite everything to a concrete, sequenced plan.
We also separated noise from signal. Some red CI runs had nothing to do with code quality: they were caused by an exhausted CI budget. We flagged that as a billing fix rather than an engineering task.
Characterization tests and safety nets before any refactor
Before refactoring anything, we added tests that lock in current behavior.
- API contract snapshot: a reflection-based test records every endpoint's HTTP verb, route template and route name (275 endpoints). Any accidental route change fails CI.
- Database model test: verifies the Entity Framework mapping (types, defaults, value generation) is unchanged, so no surprise migrations.
- Mapping regression tests: proved that request-to-entity mappings produce identical defaults. We ran them against both the old code and the new code.
- SQL byte-equivalence tests: for raw SQL query builders, we captured the exact SQL generated by the original code. The refactored builders must reproduce it character for character.
- Dependency injection resolution test: builds the real service registrations and confirms every refactored service resolves with all dependencies populated.
- Characterization tests: more than 100 tests covering every branch of complex business methods (payroll, invitations, team management, subscriptions), all passing on the original code first.
Fixing in risk-ordered, reviewable pull requests
Instead of one massive change, we shipped a sequence of focused pull requests, each verified by the quality gate on its own.
- Gate blockers first: the bugs and vulnerabilities, plus explicit request cancellation handling. This restored Reliability and Security to A.
- Mechanical cleanup: about 170 low-risk issues. That meant exceptions attached to error logs, structured logging templates, dead code removal and naming consistency. None of it changed runtime behavior.
- API surface rules: route declarations were moved to class level across 78 endpoints, and 42 request fields were made nullable with the exact previous defaults. The contract snapshot proved nothing changed for clients.
- Structural refactors: complex methods were broken into focused helpers, and oversized constructors and method signatures were replaced with parameter objects and grouped dependencies. Every step was guarded by the characterization and SQL equivalence tests.
Fixing 26 log injection vulnerabilities with one sanitizer
Rather than patching each flagged line, we added a central log sanitizer that neutralizes line breaks and control characters in user-supplied values before they reach the logger. One change protects about 260 logging call sites at once, including the ones the analyzer hadn't flagged yet.
Choosing the behavior-preserving fix, every time
Where the analyzer's default suggestion could have changed behavior, we chose an equivalent that doesn't.
- Request fields became nullable with the same fallback values, instead of being made required, which would have started rejecting requests clients send today.
- In a rate limiter, cancellation was deliberately opted out, so aborted requests don't start producing error logs.
- A hard-coded value was moved into existing resource files instead of new configuration, to avoid adding constructor parameters and deployment risk.
What the safety nets caught
A careful, test-first approach to legacy code refactoring doesn't just avoid regressions. It surfaces problems a quick fix would miss.
A dormant performance risk: a cache invalidation block had never run in production, because it looked up a member that didn't exist. Fixing it the obvious way would have switched on an expensive full-keyspace scan against production Redis. We removed the dead code, kept behavior identical, and logged real cache invalidation as its own ticket.
A fix that would have looked done but wasn't: one transactional email was actually served from a copy embedded in a resource file, not the template file the analyzer scanned. Changing only the scanned file would have turned the dashboard green without changing the real email. We updated both and added a test that loads the compiled templates.
A latent data-filtering bug: two repository methods appeared to filter on the opposite condition from their siblings. Because fixing it would change results, we documented it for a product decision instead of changing it silently.
A gap in our own tooling: the route snapshot initially mis-recorded endpoints with method-level routes. We fixed it and regenerated the baseline from the untouched code before relying on it.
Results
- Releases unblocked: the quality gate conditions for reliability and security were restored to A, and each pull request cleared its own analysis.
- Zero contract changes: all 275 endpoints unchanged, verified automatically on every build.
- Stronger safety net: more than 160 new automated tests covering the platform's most complex business logic, including areas that previously had no live test coverage.
- Security hardening across the codebase: log forging neutralized at the source for every endpoint, not just the flagged ones.
- Maintainability: the most complex methods were brought under the analyzer's cognitive complexity threshold, and constructors with 10 to 15 dependencies were reduced to 7 or fewer.
- Delivered fast with AI-native engineering: thanks to Claude Opus 5.5 and parallel multi-agent workflows, the plan went from diagnosis to a series of reviewed pull requests in under 1 week.
Before and after, at a glance:
- SonarQube quality gate: failing on every release to main, now passing on each pull request that has finished analysis.
- Reliability rating (new code): C, now A.
- Security rating (new code): B, now A.
- Open new-code issues: 376, now 0.
- Log injection vulnerabilities: 26, now 0.
- Automated tests: about 670, now 830+, with more than 160 new safety-net tests.
- Public API changes: none, all 275 endpoints kept the same verb, path and route name.
- Delivery model: from manual, issue by issue, to AI-native, with Claude Opus 5.5 orchestrating parallel agents and engineer review at every merge.
How to reduce technical debt without breaking production
Quality gates exist to protect production, but a gate that has drifted out of reach becomes a release blocker, and teams start to route around it. The fix isn't to lower the bar or mass-suppress warnings. It's disciplined engineering: measure precisely, build safety nets first, and change code in small, provable steps.
AI-native development doesn't replace that discipline, it scales it. With Claude Opus 5.5 running parallel, well-scoped agents under clear guardrails and human review, we can work through that kind of rigor far faster than a manual cleanup, without trading away safety.
FAQ
How do you reduce technical debt without breaking production?
Lock in current behavior with tests before changing any code, then fix issues in small pull requests ordered by risk. In this project, contract snapshots, SQL equivalence tests and characterization tests had to pass on the original code first, so any behavior change in a refactor failed CI immediately.
What are characterization tests?
Characterization tests record what existing code actually does today, rather than what a spec says it should do. They make legacy code refactoring safe: if a refactored method returns anything different from the original, the test fails. We wrote more than 100 of them for payroll, invitations, team management and subscriptions.
How do you fix a failing SonarQube quality gate?
Pull the issue data from the SonarQube API and separate the issues that block the gate from overall project debt. Fix bugs and vulnerabilities first, since they drive the Reliability and Security ratings, then work through lower-risk rules. Also check the new code definition: here it had silently grown to two years of code.
How do you prevent log injection in .NET?
Sanitize user-supplied values before they reach the logger, stripping or encoding line breaks and control characters, and use structured logging templates instead of string concatenation. A single central sanitizer covered about 260 logging call sites in this codebase.
What is AI-native development?
AI-native development treats AI agents as the main engineering engine rather than an autocomplete tool. Engineers set direction, define guardrails and review every merge, while agents plan, write tests and make changes in parallel.


