Skip to content

Software Engineering in the Era of AI

•

15 min read

A woman sitting on her desk confused about which AI tool to use

Software Engineering has probably changed more in the last two years than it did during the previous decade. And that's not because we discovered a new language, framework, or infrastructure model that suddenly changed everything. Nope.

The biggest change is actually more radical: we stopped writing code ourselves.

Or, to be precise, at the time I'm writing this article, most teams have started doing so (while others are still trying to figure things out).

A Change in Paradigm

Just less than two years ago AI-assisted development mostly meant GitHub Copilot guessing the next few lines of code, and software engineers were still skeptical about using it. I still remember a coworker of mine joking about times when we were on "Copilot mode" (vs "autopilot mode").

I didn't like working with AI, but I admit it was as tempting as a chocolate cake waiting in the fridge with a big label saying "DON'T EAT ME".

And somehow, in what felt like no time at all, AI tools started popping up everywhere. New editors, coding agents, model providers, review tools, testing tools, infrastructure tools - you name it. Hundreds of startups were born around AI, while established companies rushed to integrate it into products that had existed for years.

And of course, I then found myself building a real-time chat application relying heavily on an already-evolved version of GitHub Copilot, which was perfectly able to build entire features end-to-end itself. The AI tool helped me get from an idea to a working application in an afternoon, while I reviewed the output and refined the implementation, making sure the whole project looked, behaved and was set up exactly as I wanted.

Today, coding agents can inspect a repository, read specifications, modify multiple files, run tests, review their own work and iterate on failures. They even review pull requests for work described in tickets written by another coding agent, then file follow-up tickets themselves - something that shocked me completely the first time I saw it.

Engineers (like me), on the other hand, now spend their time defining the problem, providing context, reviewing decisions and orchestrating the process around the code. They can now focus on having a better product vision, focusing on UI/UX and engineering judgment of the problem.

The shocking truth is evident: what it means to be a software engineer started to change.

From Autocomplete to Agents

GitHub Copilot's technical preview launched in June 2021, describing the tool as "Your AI Pairing Programmer". It was already suggesting both lines and entire functions with just the hit of a TAB key. The interaction was simple and second nature to programmers: write some code, inspect the grey suggestion, accept or reject it, and continue.

Much faster development, and it was helping me recall methods, syntax or lint rules without much brainwork.

Then, without noticing it, we moved from auto-suggestion to have entire conversations with AI tools, expanding that interaction: you could ask for an explanation, paste an error, or discuss an implementation. But the software engineer still carried much of the context between the conversation, editor, and terminal, and nobody ever imagined (nor wanted to imagine) that a much deeper shift was about to happen.

Coding agents can now participate in that development loop. Tools such as Claude Code can inspect files, edit an implementation, run commands, and use the results to decide what to try next. They can watch for CI changes, react to a failing test, and use that feedback to push a fixed version, that will eventually either make the CI go green or the agent loop back.

In reality, autocomplete, chat, and agents still coexist today, but the distinction is in how much work we're delegating and how independently the tool can proceed. A small suggestion and a repository-wide change need very different amounts of supervision.

Specifications As Source of Truth

Take a requirement like "add a user search to the dashboard."

An agent can produce a plausible implementation quickly, but the problem is how many product decisions the AI arbitrary decided from that sentence: which users can be searched? What happens when requests finish out of order? What should someone see when the request fails?

A useful starting specification might say:

Design a responsive dashboard search bar component with real-time auto-complete, clear/reset buttons, and debounced input handling. Connect the search query state to filter an adjacent data table showing transaction history, including empty and loading feedback states. Make sure results are cached.

Those details affect the API, interface, and tests, and leaving them unspecified gives the agent room to choose on your behalf - which is not always what we want.

Of course, the first specification will be incomplete. Building the feature may reveal an assumption that needs changing. When that happens, the requirements and affected tests need updating alongside the code. Otherwise, you've just created three different versions of what the feature is supposed to do.

This is why writing a specification is engineering work. It exposes decisions that a vague ticket hides. It also gives you something concrete to review the implementation against.

There's a whole engineering methodology called Specification Driven Development (this other article is a great one too) that defines how to build software using specifications and guardrails in order to guide AI agents towards the final product.

Your Codebase Is Now Part of the Prompt

An agent needs to discover conventions that a teammate might already know: where business logic belongs, which commands run the checks, and why a particular abstraction exists.

The good news is that it does it automatically like a human might do: inspecting the codebase, checking the code around it, "copy-pasting" from a similar component. That's great, but it burns too many tokens.

That's why files such as CLAUDE.md, .cursorrules or the increasingly adopted open standard AGENTS.md provide a place for that context. These files are an entry point for coding agents so they can be instructed on which folders to read, how to run the project or which coding styles we prefer over others.

But creating an enormous instruction file containing every fact about your codebase isn't the solution either, as this also *burns tokens fast.*

A better approach is treating the root instructions as an index: explain how the project works, where deeper documentation lives, how to run validation and which constraints agents must respect.

A modern AI-friendly repository might contain something like:

bash
AGENTS.md
ARCHITECTURE.md

docs/
├── architecture/
├── decisions/
├── specs/
├── plans/
|
skills/
└── ...

The loading conditions also matter - "Architecture notes live in this folder" is less useful and creates more context load than "read these notes when making changes to the API layer", as with the first we implicitly say to load the whole folder each time considering any architectural decision.

Documentation also needs to agree with the repository: if the instructions describe a pattern the code abandoned six months ago, the agent has conflicting examples to follow.

DORA's 2025 research describes AI as amplifying an organization's existing strengths and weaknesses. To me, this is a practical reason to invest in clear boundaries, reliable checks, and documentation someone can actually maintain.

Reviews Are More Important Than Ever

In Stack Overflow's 2025 Developer Survey, 46% of respondents to the accuracy question said they don’t trust AI tools, compared with 33% who trusted them. That doesn't measure how often generated code is wrong, but it does show that using these tools and trusting their output are different things.

And I have experienced this myself: in my current team we're developing software using specification driven development, and even though we have dozens of .md files, guardrails, a very optimized CI/CD pipeline and tests written for almost everything, the AI is still failing.

Do not misunderstand me: AI does a lot, it does it fast, and at an incredibly high engineering level, but engineering decisions are still on the person who writes the prompts. Most importantly, the output needs to be accurately tested on a running environment, and often times there are edge cases the AI fails to spot, so can’t be trusted blindly - at the same level we never trusted blindly neither our most senior engineer in the team!

Getting back to the search example, an agent might write an implementation and tests that both assume every user is visible to every account, but that might not be what we wanted. Validation errors may push down the layout or a particular banner component may have a background very different from what's everywhere in the app... perhaps only because we've chosen a cheaper AI model to run our prompt through.

The more tests and guardrails we write, the more AI agents write the perfect feature - still, in my opinion, engineering judgment and human steering are mandatory, important parts of the whole development process, and perhaps more important than ever.

What Orchestration Actually Involves

In practice, orchestration means deciding which work to delegate, what information it needs, and how to evaluate the result. It can be as simple as one agent working through a well-defined task.

Before handing a task to an agent, I'd want to establish:

  • Scope: what it should change and which decisions need clarification.

  • Context: the relevant requirements, code, and project conventions.

  • Verification: what evidence would show the work is complete.

  • Dependencies: what must happen first and what can proceed independently.

That last point matters when using several agents. Two independent investigations may run well in parallel. Two agents changing the same shared module can leave you with an integration problem. Someone still needs to reconcile their assumptions and review the combined result.

Agents are also becoming increasingly connected to the systems around them. Through standards such as the Model Context Protocol (MCP), they can be given structured access to tools and external data sources - from documentation and databases to issue trackers and internal APIs - instead of relying only on what exists inside the repository.

Of course, giving an agent more tools also means giving it more permissions, so deciding what it can read, modify or execute becomes another engineering decision.

Model choice is also important: in my own workflow, I constantly pay attention to reasoning effort and token consumption, as tokens depend on the subscription model an account follows. If an agent repeatedly heads in the wrong direction, I'd revisit the task and its context before spending more on another attempt.

Behind all these coding agents are Large Language Models (LLMs), and choosing which one should handle a task is becoming part of the engineering workflow itself. Different models behave differently when reasoning about architecture, generating code, following long specifications or working autonomously, so using the most powerful model for every task isn't necessarily the best choice.

The effort is the other variable at play: they tell coding agents how strongly they should think about a problem, how much to iterate, how many edge cases to consider, and so on. Higher is not always better, and not only for a context / token usage issue, but because the agent may over-engineer around the problem, which is not always what we want.

A common AI Engineering Setup
A common AI Engineering Setup

Technical Interviews Are Changing Too

In my article about debugging and refactor a real product component, I wrote about an exercise that reveals how candidates reason through realistic problems, and how I believe it's one of the best technical challenges - at least at the time of writing - from both an interviewer's and candidate's point of view.

With AI starting to be more and more integrated in today's interview processes, the mental model of the problem, architectural decisions and performance bottlenecks to solve are still valid: candidates are still expected to know which engineering problem they need to solve, what are the trade-offs and what is the best way to do so. However, how they should approach the challenge is in a completely different dimension.

I have taken full screening interviews conducted by AI that was reacting and asking follow-up questions based on my responses, and that was already shocking. I still haven't, however, participated or conducted an AI technical interview myself, but I'm well aware they are very different from the ones we're used to.

In some newer interview formats, candidates are expected to be less focused on recalling syntax and more on working alongside an AI agent. The challenge may be to work on an existing component generated by AI or to set up a full project environment according to the given task specifications, and tell the interviewer how they would interact with the AI agent.

In either case, do they clarify the requirements? Are they noticing an unnecessary dependency? Are they checking the behavior, or stop when the tool says it's done? Equally importantly: are they managing context properly and knowing which LLM or agent would be best suited for the task?

CodeSignal already offers AI-assisted interviews, so this is more than a hypothetical format.

Today’s problem is that you may interview with one company that prohibits AI completely and expects you to solve an algorithm on a blank editor, while the next may explicitly ask you to use Claude Code, Cursor or Copilot to ship production software, and judge if you tremble on doing so.

Framework expertise are gradually losing some weight, where engineering fundamentals aren't: architecture, debugging, security, performance, data modelling and the ability to understand unfamiliar systems arguably become more important when you're reviewing code you didn't personally write.

That difference says quite a lot about the engineering culture you're joining, and how AI integration in software engineering is still taking place. But it's happening, and it's happening fast.

The Bottleneck Just Moved Up

There's another consequence of being able to generate code so quickly: there's now so much code to review and PRs open that the bottleneck now moved one step up in the pipeline, at review and manual testing level.

In short: someone still needs to decide whether that code should actually ship.

Review and QA are incresingly becoming the bottleneck of a development pipeline
Review and QA are incresingly becoming the bottleneck of a development pipeline

When writing an implementation took days, spending a few hours reviewing it wasn't particularly problematic. But if an engineer can have several agents opening pull requests in parallel while they're working on something else, suddenly the production rate becomes much higher than the human review capacity.

And that's exactly what some teams are already seeing. CloudBees' 2026 research found more engineering leaders pointing at reviewing, testing and deploying code as their current delivery bottleneck than at writing the code itself.

PostHog is an interesting example of this transition. Agents now open around 70% of the pull requests in their monorepo, and the company has openly described human review as something that quickly becomes impossible to scale when engineers are dealing with dozens of agent-generated changes. Their answer wasn't simply review faster, but to introduce more automation around review and let agents handle low-risk approvals while humans remain involved where judgment actually matters.

But here's where things become even more interesting: some teams are removing human code review almost entirely.

In OpenAI's experiment building a product with Codex, pull requests can be reviewed by other agents, iterated on automatically and eventually merged without a human having to approve every diff. OpenAI actually reported that, once code generation and review scaled, their bottleneck moved again - this time to human QA capacity. Engineers were no longer struggling to produce or review code: they were struggling to validate all the resulting behavior.

StrongDM's AI team has taken an even more radical approach with what they call a Software Factory. Their rules explicitly state that code must neither be written nor reviewed by humans. Instead, humans define intent, constraints and scenarios, while agents generate the implementation and repeatedly validate it against automated behavioral checks until it converges. In that setup, validation replaces code review.

This doesn't mean that removing human review is suddenly a good idea for every codebase.

It works much better when success can be clearly observed: an API returns the expected result, a workflow completes correctly, performance stays below a known threshold, or an interface behaves according to a well-defined scenario. If the real question is whether an abstraction will still make sense two years from now, whether the UX actually feels right, or whether a requirement itself is wrong, it's much harder to encode that judgment into a test.

And different environments simply move the bottleneck somewhere else. A team with weak tests may become constrained by QA. A legacy codebase with poor documentation may hit a context bottleneck, because agents spend too much time understanding how the system works. Highly regulated teams may hit security and governance limits. A mature agent-first team might automate all of those and eventually discover that product decisions, specifications and human attention are now the scarce resources.

From Coders to Builders

For a long time, being a software developer was closely associated with the ability to write code.

Of course, senior engineers have always done much more than that: architecture, product discussions, debugging, mentoring, trade-offs and reviewing have always been part of the profession.

Today, what is shifting is the balance: engineers increasingly operate one abstraction level above the implementation itself.

And I lived that personally: instead of spending an afternoon manually creating API handlers, database models, UI components and tests, I can now start working with something closer to:

This is the feature. These are the constraints. This is how it should behave. These are the existing patterns you need to follow. Build it, test it, and show me what changed.

And then orchestrate, steer, review, take product-level and UI/UX decisions, and.... basically, build instead of coding.

Full-stack awareness has become more valuable because the cost of crossing traditional boundaries is lower: a developer who historically specialized in one area can now navigate unfamiliar parts of a system much faster with the help of an agent.

Product intuition becomes more valuable because producing code is not the bottleneck anymore, so we can now move one level of abstraction higher and do more.

My role has moved from coder toward builder and, you know what? Although it wasn't easy to accept this change at first, it is so damn addictive, and as rewarding and fulfilling as ever.

Personal Conclusions

Teams are adopting these tools at different speeds, and even if the AI era is already here and is evolving fast, it seems like not everyone feels ready. There's still a lot of confusion around: companies are betting on AI, devs still doubt, and I personally know quite a few colleagues who moved away from software engineering to dedicate themselves to completely different professions and activities, like woodworking or fashion design.

I had my existential crisis too, not so long ago, and I'm so glad to have overcome it and still feel part of the game. I actually feel like I'm growing and, generally speaking, more fulfilled than ever, seeing my skills evolving one step higher and towards new limits.

And - to conclude - even though the rise of AI-assisted coding is evident, and feels very similar to how we moved from writing ASM to using high-level languages, I still can't be sure we're moving towards a completely codeless paradigm. So, keep your skills sharp.