The Skills They Called Second-Class Run Every AI Agent Now
The QA engineer skills for AI engineering are not a new curriculum you need to buy. They are the muscle the industry mocked for a decade: frame the work precisely, then judge whether the output survives production. AI moved the executor inside the frame from humans to agents, and the human job collapsed into exactly two things. Framing and validation. That was always the QA job. Now it is everyone's job.
Why was QA branded second-class engineering in the first place?
It was branded second-class for the wrong reasons, by people who confused who writes the code with who owns the system. The label said: you don't build, so you don't matter. The caste was clean. Devs created, QA reacted. One side made things, the other side filed tickets that ruined someone's Friday.
I spent years outrunning that label. Not arguing with it, outrunning it, which is what you do when the room has already decided what you are. You take the boring incident nobody wants and you turn it into a system. At Stenn I invented Bug-mits, a weekly cross-team incident forum, strictly blameless, and production criticals dropped 33 percent. That was one factor inside a broader quality policy, not a magic number, but the direction was unambiguous. I ran the same system at a crypto company and production criticals fell more than 50 percent, with ownership moving onto the teams instead of living in my inbox.
Here is the irony I sat with for years. The work that got me dismissed as a non-builder was the work that scaled the whole org. Fixing the class, not the incident. A $1.5M wrong charge gets handled the same calm way as a broken CSS rule: stabilize it, then lock the fix in writing with a named owner so the class retires forever. That is not reacting. That is governing a system. The people who called it second-class were standing next to the most leveraged seat in the building and reading the nameplate wrong.
How did AI make framing and validation the entire engineering job?
AI became the execution layer. That is the whole story, and most people are narrating it as a revolution when it is a substitution. The discipline did not change. Delivery management was always about making framing and validation manageable, scalable, predictable. AI changed who executes, humans out, agents in. The job description stayed.
Watch what is left when an agent writes the code. You no longer type the implementation. You write the spec sharp enough that a tireless executor can't misread it, you point it at the problem, and then you stand at the gate and decide: does this survive contact with production, or do we throw it out. Two human acts. Input framing and output validation. Everything in the middle is autonomous by design.
The caste that separated dev from QA is dying from the bottom up, because the engineer's actual daily motion is now framing a request and validating what comes back. The dev just became a QA engineer with a fleet to run.
I built an operating model around this, managed autonomy: 7 stages, 2 human gates. The middle is the agent's territory, untouched. The two gates are human because that is where leverage concentrates. Input, where you frame. Output, where you judge. Now read those two words again. Framing. Judging. That is the exact shape of QA work, promoted from a corner of the org to the entire engineering loop. The caste that separated dev from QA is dying from the bottom up, because the engineer's actual daily motion is now framing a request and validating what comes back. The dev just became a QA engineer with a fleet to run.

Why do QA engineers have an edge with AI agents?
Because the agent fails in exactly the ways QA people were trained to anticipate, and most builders were trained to ignore. A QA engineer doesn't read code and feel proud. A QA engineer reads code and asks where it breaks under load, what the empty state does, which assumption is going to detonate at 2 a.m. That is the systems-judge reflex: not builder overclaim, but a flat 'this won't survive production, rewrite it.' Agents need that judge constantly, because they are confident in proportion to nothing.
There is a phrase floating around that AI output is bad. It is a skill issue. When someone shows me garbage from an agent, the garbage is almost never the agent. It is the framing. Vague prompt in, vague system out. The people complaining loudest about hallucinations are usually the ones who handed the model a sentence and expected an architecture. QA people don't do that, because we have spent careers writing test conditions that leave no room for interpretation. Precise input is our native language. We were writing prompts before prompts had a name.
And we own the reflection loop without being told. Every time an agent returns something broken, the QA instinct is not to patch the output, it is to fix the input that produced it, so the entire class of defect retires. Builders fix the bug in front of them. Judges fix the frame that generated the bug. With a fleet of agents, the second behavior is the only one that scales.
The clay anti-pattern: why most engineers eat AI from the tail
Most engineers are eating the clay from the tail. Picture a strip of clay. You can pull it from the head, the prepared end, where it comes off clean and continuous. Or you grab the tail and rip, and it tears and crumbles and you spend the afternoon picking pieces off the floor. Eating from the tail means: skip the framing, fire a lazy prompt, then spend three hours wrestling the output into something usable. The tail feels faster for the first ten minutes. It is the slowest possible path.
The law of preparation is brutal and simple: prep quality drives output stability. There is no second lever. I once generated 1300 articles through a pipeline, and the entire game was upstream. Get the preparation right and the volume just falls out the other end. Get it sloppy and 1300 becomes 1300 separate fires. The output was never the work. The framing was the work, and the output was a consequence.
The people who branded you are still optimizing for typing speed in a building where nobody types anymore. The skills they called second-class for AI engineering are not a side track. They are the job description. You just got promoted, and your competition is still arguing with the nameplate.
This is where the old caste collapses for real. The builder mindset says the value is in the generation, so the builder rushes to generation. The QA mindset says the value is in defining what correct means before a single line exists, so it lives at the head of the clay by default. When the executor is an agent that will happily generate ten thousand confident wrong lines, the head of the clay is the only seat that matters. The people who learned to skip preparation because they were fast typists have lost their advantage. Typing speed is worth nothing when you are not the one typing.
What does a senior engineer own when agents write the code?
A senior engineer owns the frame and the gate, and is accountable for both in writing. Not the keystrokes. When agents write the code, the seniority moves entirely into judgment: what to build, what 'done' means, and whether the result is allowed near customers. You become the conductor of the fleet, and conducting is not playing every instrument, it is deciding what the music is and stopping it when it goes wrong.
Concretely, three things sit on your desk. First, the spec sharp enough that an agent cannot drift, because ambiguity is now a production risk, not a style preference. Second, the validation bar, the explicit definition of survival, because an agent will declare victory on garbage and mean it. Third, the reflection discipline, where every failure returns to the input and retires a class instead of a ticket. Process becomes the governor here. When leverage is this high, a single bad frame ships at machine speed across the whole fleet, so the governor is not bureaucracy, it is the thing that keeps the engine from flying apart.
I learned what skipping the frame costs the slow way. On Dream Book I spent 4 months blaming a 50 percent drop-off on 'people don't remember dreams' before I bothered to look at onboarding. The problem was never the users. It was a frame I never validated, an assumption I let run unexamined. The fix was upstream, where it always is. Contrast that with Dating Coach, where I rewrote the titles and CTR went from 0.4 percent to 6 percent. Same lesson, opposite outcome: when you actually frame the input precisely, the output moves an order of magnitude. The senior job is owning that frame, and owning it out loud.
Stop outrunning the label. The label is the job now.
For years I treated QA as something to escape. Situation: branded second-class, watched the room hand respect to whoever shipped the most lines. The real pain was not the dismissal. It was building the most leveraged system in the org and still being read as support staff. The durable lesson is narrow and I will keep it: the work was never lesser, it was earlier. Framing and validation always sat upstream of the code, which is exactly why it was invisible to people who only counted output.
We are already in cyberpunk and most engineers haven't clocked it. The executor is now a fleet of tireless machines, and the only human acts left are deciding what to build and judging what came back. That is the QA muscle, with the whole org on top of it. So stop outrunning the label. The people who branded you are still optimizing for typing speed in a building where nobody types anymore. The skills they called second-class for AI engineering are not a side track. They are the job description. You just got promoted, and your competition is still arguing with the nameplate.