Yesterday evening, a teammate opened a huge spreadsheet that had been created for a client.
It looked substantial. It looked organized. It looked finished.
The client had already found something important that it failed to address.
My teammate’s private message to me was blunt:
“AI slop. Zero human review.”
Harsh. But true.
The criticism was aimed at another teammate. I was the one who turned it back toward me.
“I might be the worst offender,” I admitted.
They did not exactly disagree.
They said they do review some of my work before presenting it to clients, but that this was “a known Mark factor.”
Fair. 😅
The joke contained an uncomfortable truth. I had recently responded directly to a client using AI output I had not properly reviewed. Like the spreadsheet, it looked finished. It sounded confident. I assumed it was good enough.
It was not.
The client was pissed off. Rightfully.
That is the deceptive thing about AI-generated work. The failure is not always obvious. The output can be polished, thorough, and professional while still missing the central point.
We are accustomed to unfinished thinking looking unfinished. AI breaks that relationship.
The obvious lesson is that AI output needs human review. But that explanation feels incomplete. We all know these systems make mistakes. We have repeated “human in the loop” so often that it has become almost meaningless.
The more difficult question is why intelligent people who understand these systems keep sending their output into the world without fully understanding it.
I think part of the answer is that AI has changed the economics of producing plausible work without changing the economics of judgment.
Production became cheap
Before generative AI, producing a detailed analysis, client response, or nine-tab spreadsheet imposed natural limits.
The work required enough time that the person creating it usually developed some understanding of it along the way. Writing was not separate from thinking. The friction of production forced a certain amount of contact with the material.
AI weakens that connection.
We can now produce the visible evidence of thought much faster than we can perform the thought itself. A coherent document can appear in minutes, complete with headings, recommendations, caveats, and a confident conclusion.
It looks like the end of a process, even when it is closer to the beginning.
That surface quality matters because we often use polish as a proxy for completeness. AI output does not arrive looking like a rough draft. It can look better than many human final drafts.
But polish is not judgment.
A document can be clear, professional, and completely misunderstand the assignment. A spreadsheet can contain nine beautifully organized tabs and still fail to answer the client’s actual question. A response can sound empathetic while missing the one fact that matters.
The problem is not merely that AI can be wrong. People are wrong all the time.
The problem is that AI can be wrong in a form that looks unusually complete.
The pressure is coming from both directions
None of this is happening in isolation.
Our clients know these tools exist. Many are using them themselves. They reasonably expect us to move faster, operate more efficiently, and accomplish things that would have required more time and larger budgets a few years ago.
The pressure is also coming from inside the company.
I am actively pushing our team to use more AI. I want us automating repetitive work, researching more broadly, testing more thoroughly, and finding ways to deliver better results without simply adding more hours.
I believe that is the right direction. Refusing to use these tools would not preserve some higher standard of craftsmanship. It would make us slower, more expensive, and eventually less useful to our clients.
But I also have to acknowledge the tension I am creating.
I am asking the team to increase its output while expecting the same people to preserve the judgment, context, and attention that made the work valuable in the first place. Clients are raising the bar from the outside, and I am pushing for greater efficiency from the inside.
The team is caught between those expectations.
When something goes wrong, it is easy to point at the individual and say, “Zero human review.” Sometimes that criticism is deserved. People remain responsible for the work they send.
But if leadership rewards speed, expands the amount of work people are expected to handle, and immediately fills every hour AI saves, inadequate review is not only an individual failure. It may also be the predictable result of the system we created.
AI can reduce the cost of production. It does not eliminate the cost of verification.
If I want the team to use more AI, I also need to make room for the work that follows: checking assumptions, tracing conclusions, questioning whether the output answers the actual question, and deciding whether we are willing to put our name on it.
We cannot demand acceleration and treat review as free.
Clients increasingly expect both speed and quality. Internally, I expect us to use the best tools available. But if every efficiency gain becomes an excuse to take on more work, then the time saved by AI never becomes time available for judgment.
The bottleneck has not disappeared.
It has moved from producing the work to understanding it.
Every output creates review debt
I have started thinking about this as review debt.
Every AI-generated artifact creates an obligation to verify it. The faster we produce artifacts, the faster those obligations accumulate.
Like technical debt, review debt can remain invisible for a while. The documents exist. The tasks are marked complete. The client receives something on time. From the outside, the system appears more productive.
But the uncertainty has not disappeared.
Someone still needs to confirm that the analysis used the right assumptions, that the spreadsheet answers the actual question, and that the client response reflects what really happened.
If that work is not performed before delivery, the debt remains embedded in the output. Eventually, someone pays it through rework, confusion, damaged trust, or an uncomfortable client call.
Sometimes the work has not even been eliminated. It has simply been displaced.
When my team quietly reviews and repairs my AI-assisted work, I may experience a productivity gain that does not exist at the company level. My time was saved by consuming someone else’s attention, often later in the process and under greater pressure.
The work moved from the visible author to an invisible reviewer.
Calling that efficiency would be misleading.
Hierarchy makes the problem worse
There is another uncomfortable dimension to my experience: I am the CEO.
People are less likely to tell me directly that my work looks like AI slop. They may revise it, compensate for it, or develop an affectionate term like “the Mark factor.”
Hierarchy makes direct feedback harder. The people with the most authority may receive the least correction.
That is especially dangerous with AI because confidence scales more easily than competence.
AI allows me to involve myself in more subjects, generate more opinions, and produce more material. Each output carries the implied authority of the person sending it, even when that person has spent very little time developing the underlying judgment.
The result can look like expanded leadership capacity while functioning as expanded organizational noise.
That possibility bothers me more than the embarrassing client response.
A single bad message can be corrected. A company quietly learning to compensate for its leader is a structural problem.
I tried to solve judgment with more engineering
I also blamed the models.
As I moved between newer models, I noticed the output becoming more verbose and changing in ways I had not fully anticipated.
I eventually built a separate skill using GPT-5.6 to polish the output from another model.
There is something revealing about that response.
I had one system generating the work and another system improving its presentation. The final result read better, but the additional polish made it no more likely that I understood or agreed with the substance.
I was improving the signal that told my brain the work was finished.
The actual problem was not verbosity. It was that I had substituted a chain of increasingly polished outputs for my own judgment.
Better prompting can improve an answer. A second model can identify errors in the first. Automation can create valuable checks.
But none of those things resolves the question of who understands the work well enough to take responsibility for it.
“Human in the loop” is not a process
It is easy to conclude that a human should remain in the loop.
The phrase sounds reassuring but leaves the important questions unanswered.
Which human? At what point? With how much time?
Are they reviewing the presentation, checking the facts, or reconstructing the reasoning?
Do they have the authority to send the work back when the deadline is approaching?
A quick glance from an overwhelmed person does not become quality control merely because a human technically participated.
Real review has a cost. If we want it, we have to account for it when setting timelines, assigning ownership, and deciding how much work AI allows us to take on.
The standard cannot simply be that a person looked at the output.
Someone must understand it well enough to explain the reasoning, defend the conclusion, and accept responsibility when it is wrong.
The goal is not maximum output
AI is not making human judgment obsolete. It is making judgment the scarce resource.
That changes how I think about productivity.
The goal should not be to maximize the volume of work we can generate. It should be to maximize the volume of work we can responsibly stand behind.
Those are not the same thing.
The companies that use AI well may not be the ones producing the most. They may be the ones that learn where speed is valuable, where friction is protective, and where human attention cannot be removed without changing the nature of the work.
We went through a similar process with remote work. The technology made distributed work possible before companies understood the norms required to make it sustainable. Over time, we learned that trust, accountability, asynchronous communication, and real-time collaboration each had their place.
We are still developing those norms for AI.
For now, I know I need to change my own behavior.
AI can help me begin the work. It can challenge it, organize it, and test it. But if my name is on the output, understanding it is still my responsibility.
AI did not eliminate the final round.
It made the final round more important.


