hero image

The Content Factory Works Because No Single Model Gets to Be the Hero

July 22, 2026

The Content Factory Works Because No Single Model Gets to Be the Hero

I stopped asking one AI to do everything and started building a small assembly line instead.

I get asked a version of the same question all the time.

What model are you using now?

People want one name. One tab to keep open. They want a single subscription to justify. One answer they can paste into their own workflow and feel done.

That answer would be nice. It would also be wrong.

I want to show you the inside of something I have been running quietly for a while now. Not to impress you. Honestly, parts of it are a bit embarrassing in how cobbled together they look. But it works. The full cost stays under $1 per finished article.

The contrarian truth is simple. The model everyone calls “the best” right now is actually the worst choice if you use it for everything.

Every flagship model is a generalist. It is good at many things. It is great at none of them in isolation. When you route everything through a single model, you inherit all its tics.

GPT-whatever writes confidently but has a breezy sameness across paragraphs. Claude is more careful and interesting sentence to sentence, but sometimes flattens an argument. Gemini reasons well across long documents, but the drafts it produces read a little clinical.

So I do not pick one. I pick each of them for the job they are actually good at.

Not because the specialist chain is clever. Not because I enjoy making things more complicated than they need to be. I do not. I like boring systems. I like things that survive Tuesday afternoon when I am tired and half paying attention.

I built this factory because one model never wins on every metric at once.

So I stopped asking for a winner.

I started assigning jobs.

The Total Pipeline

That is the real story behind the content pipeline I am using now. It writes the article, creates images, makes social variants, spits out a vertical video, and gets from prompt to published post in under fifteen minutes on a normal run.

That line matters to me.

Not because I am trying to save 14 cents like a maniac. I am not. It matters because once content gets cheap enough, the real bottleneck becomes taste. The real scarcity is judgment. You have to know when a paragraph sounds dead even though it is technically fine.

Step One: Parallel Drafting

The stack starts with parallel drafting.

I send the same brief to GPT-5.5 and Claude Sonnet 4.6 simultaneously. Same guardrails. Same target reader. They both produce a full draft.

Why two?

Because they fail differently.

GPT-5.5 tends to nail structure and flow but goes generic in the examples. Claude Sonnet 4.6 writes more distinctive sentences but sometimes buries the lead or goes long on a point that did not need the space. One drifts into polished sludge when it wants to impress me. The other gets stiff when it tries too hard to behave.

If I only use one, I inherit one model’s blind spots and call it style.

That is a bad trade.

Two drafts in parallel cost a bit more up front. They save me time where it counts.

Running them in parallel rather than sequentially matters. If I ran GPT first and fed that output to Claude, Claude would just be editing. I would get GPT’s structure with Claude’s polish. Running them in parallel means I get two genuinely independent takes on the same brief. The differences between them are useful information. Where they agree, I trust it. Where they diverge, I have choices.

Editing from abundance is easier than dragging a single weak draft uphill. I am not staring at an empty page.

Step Two: The Judge

Then comes the part people miss.

I do not ask one of those drafting models to polish itself into greatness. I hand both drafts to Gemini 3 Pro and use it as a judge.

This is a different job.

Gemini 3 Pro sees both drafts and a scoring rubric. It does not write from scratch. It picks.

Writing and picking are not the same skill. We all know this as marketers. The person who writes ten hooks is not always the person who can spot the one hook that will pull. Same here.

Gemini’s job here is to identify which draft handles each section better and stitch together the stronger version, paragraph by paragraph if needed. It flags anything that needs a fresh pass. It cuts overlap. It fixes weak transitions. It returns one merged piece.

I tested running just Gemini for the whole thing, start to finish. The output was fine.

But fine is the ceiling when you ask one model to do the full job. When it is judging instead of originating, the ceiling goes up because it works with better raw material than it would produce itself.

This part surprises most people. They expect the expensive step to be the drafting. It is not. The judge pass is where the actual quality decision gets made.

Late-night editing: two draft variants annotated by hand under a desk lamp

Step Three: The Sanitizer

Before anything goes to a human for review, the merged draft runs through a regex-based sanitizer and rewrite loop.

This is the least glamorous part of the chain. It is also one of the most useful.

A lot of bad AI writing is already a solved problem. I do not need a genius model to tell me that certain phrases sound fake.

The sanitizer has a kill list. These are words and phrases that no model can be trusted to avoid on its own. Banned phrases include “delve”, “in today’s rapidly”, “unlock”, “empower”, passive constructions in the first sentence, and fake casual openers like “let’s dive in”. The list is around 60 entries right now.

When the sanitizer hits one, it flags it.

This is a case where a regex loop beats an LLM prompt. Humans already solved this problem. We know exactly what bad AI writing looks like. The regex does not hallucinate a replacement or decide the phrase is actually fine this time. It catches it every single time. It has zero creativity.

Then the flagged text goes back through a rewrite loop.

The rewrite loop is GPT-5.5 again, but now it works sentence by sentence on a specific fix. Smaller context window. Tighter instruction. Cheaper per token. Keep the point. Keep the facts. Keep the voice. Rewrite only the flagged lines.

That matters.

If you ask a model to make it better, it will often repaint the whole house when all you needed was to replace one broken hinge. I do not want broad creativity here. I want controlled cleanup.

And yes, I still review it.

I am not pretending the machine has taste and I have retired to some villa. I still cut lines. I still swap examples when they feel too vague. But I am editing a draft that already has shape and decent manners.

That is a very different task from writing from zero.

Step Four: Images

Images run through the same logic.

Again, I do not look for one winner.

Two models handle image generation in parallel. I use GPT Image 2 and Seedream 5. They do not do the same job stylistically. I route different image types to each depending on what the article needs. Editorial headers go to one. Supporting diagrams or lifestyle frames go to the other.

One is often better at following the brief. The other sometimes gives me a more interesting texture. I am not loyal to either. I want an image that looks intentional, not like the usual AI soup with suspicious fingers and text that mutates into runes.

Generation is only half of this step. The other half is QA.

I run a vision LLM image QA. I do not review the images myself at this stage. A vision model does the first pass with a scoring prompt. It checks for obvious tells. Hands with the wrong number of fingers. Garbled text. Facial geometry that looks wrong. Off brand layouts.

This part has saved me from embarrassing publishes.

Humans are bad at reviewing the tenth AI image in a row. By image eight, your brain fills in missing details and starts grading on a curve. The vision model does not get bored. It catches little uncanny tells that slip past me on a fast review.

Images that fail the score threshold get regenerated automatically. Images that pass get flagged for a human glance. At that point, it takes about fifteen seconds.

A lot of people treat image QA as overkill. I think the opposite. Once generation gets cheap, checking gets more important.

Kitchen mise en place: four distinct knives laid out on a counter — the right tool for each cut

Step Five: Video

Every post gets a vertical video variant for social.

For this, I use Veo 3.1 Fast.

I am not going to oversell what it produces. It is not cinematic. Most of the time it does not need to be. It is appropriate. It is fast. It feels native enough to post.

The prompt for the video is generated from the article summary, not written manually. That keeps the video conceptually connected to the piece without me having to think about it separately for each post.

Fast matters here.

If the video step turns a 15-minute pipeline into a 45-minute one, I stop using it. Simple as that. The factory has to survive contact with reality. Acceptable output now beats theoretically better output later.

People get this backward with AI content. They chase peak quality in one isolated step and ignore the total flow.

I care about total flow.

If a slightly weaker image model gets me there in a quarter of the time, and the QA step catches the misses, that combo wins in practice. Not on Twitter. In practice.

The True Costs

The cost side is less dramatic than people expect.

A full run creates one 1,200-word article, three to five images that passed QA, and one vertical video.

The total is under $1. Usually around $0.73 to $0.89 depending on how many image regeneration cycles the QA step triggers.

If you want a line you can check against your own API bills, GPT-5.5 currently costs roughly $0.15 per 1,000 output tokens. A 1,200-word article is about 1,600 tokens of output. Two parallel drafts plus a sanitizer rewrite loop comes to maybe $0.12 to $0.18 of that total. The images and video are where the rest goes.

The time from prompt to a published post with social variants is under fifteen minutes. Not because any individual step is fast, but because the steps are running in parallel. The human review points are minimal because the QA is already done before anything lands in front of me.

That does not mean I publish everything untouched in fifteen minutes.

It means the machine gets me to the part where my brain is useful, fast.

That is the whole point. I do not want AI to replace judgment. I want it to stop wasting judgment on tasks that should already be mechanized.

Assembly Line Complexity

I want to be straight with you about this. The chain I just described is not simple to build.

There are about a dozen places where something can break. A model returns a malformed response. The vision QA scores something incorrectly. The video prompt generates something off topic. I have spent more than a few evenings debugging one handoff or another.

But this is not complexity for show.

It is assembly line complexity.

The same reason a decent kitchen has different knives. The same reason a media buyer uses different views before making a decision. Different jobs punish different weaknesses.

Why run it?

Because the alternative is worse. If I ran one model doing everything, I would get consistent output at a lower setup cost. But the ceiling on quality is lower. The floor on slop is higher. You spend less per piece than a freelancer. The fifteen-minute turnaround means you can react to what is ranking without a week of lead time.

There is another reason I like this model chain. It makes failures easier to diagnose.

If a final article feels flat, I can ask where it went wrong. Was the brief weak? Did GPT-5.5 and Claude Sonnet 4.6 misread the angle? Did Gemini 3 Pro merge too cautiously? Did the sanitizer loop strip too much personality out of a section that was fine?

When one giant model does everything, every problem gets blurred into one complaint. “The AI was not good.”

That tells me nothing.

A chain gives me handles.

I can swap a step without rebuilding the whole factory. The chain will keep changing. Models update. Pricing shifts. If a cheaper image model gets good enough, I drop it in. The specific tools in here are less important than the logic behind them. The system is stable because the roles are stable, even when vendors change.

These are tools with job descriptions. Nothing more noble than that.

If you write your own content, the main thing a setup like this buys you is not volume. Volume is the least interesting benefit. What it really buys you is energy.

You stop spending your good attention on first-draft sludge. You spend it on the parts readers actually feel. The sharper example. The sentence that sounds like a person said it because a person did.

One last thing.

A lot of marketers still ask AI to be a genius intern.

That is the wrong mental model.

I think of it more like a small team of uneven specialists. Some are fast. Some are cheap. All of them are occasionally weird. One writes a strong opening and a bad middle. Another organizes well. A third spots junk. One catches visual mistakes I am too tired to see.

None of them should run the whole shop alone.

That is why the factory works.

Not because I found the perfect model.

Because I stopped looking for one.

Karl G
Karl G|admin
Karl G Olsson
Back to Blog