Every content platform now has an AI feature. Sanity has AI Assist, Webflow has an assistant in the canvas, WordPress has a dozen plugins competing for the same space, and every one of them demos beautifully: a cursor blinks, a paragraph appears, everyone nods. Then you go back to your actual job, which was never really about writing one paragraph. It was about forty collection items, a price that is wrong on nine pages, and a metadata pass nobody has done since launch.
So it is worth being precise about what you are evaluating. Below are the seven things that, in our experience building and running these systems, separate a writing aid from an assistant that can genuinely operate a website.
1. Does it read your content model, or just your prose?
This is the first and biggest divide. An assistant that only sees the field you are typing in can help you phrase things. An assistant that reads your schema knows that a blog post requires a slug, an excerpt, a category reference, a reading time and an author, and can therefore produce a document that validates on the first attempt.
The practical test: ask it to create a new entry of one of your custom types. If what comes back is body copy you then have to fit into fields yourself, it is a writing aid.
2. Can it work across many entries at once?
Most of the painful work in a CMS is repetitive rather than creative. Changing a product name across three hundred entries, adding meta descriptions to every page missing one, updating a phone number in every footer variant. An assistant confined to the document you have open cannot touch any of it.
If it can do bulk work, then the follow-up question matters more: does it show you the scope first? You want to be told how many entries are affected and shown a sample of the exact change before it runs — and you want the whole operation recorded as one reversible action, not three hundred separate edits.
3. Does it respect draft and publish as separate things?
Editing was never the dangerous act. Going live was. A serious assistant treats those as two different operations: work lands as a draft by default and publishing is an explicit, separately permissioned step.
Watch for the subtle version of this. Editing a page that is already published is publishing, in effect — so any write that changes live content immediately should require publish permission too. Tools that get this wrong feel safe right up until the afternoon they are not.
4. Are permissions enforced, or merely described?
Ask where permissions live. If the answer is that the assistant has been instructed to be careful, that is not a permission system — instructions are persuadable. What you want is enforcement at the connection level, checked before each write executes, so a denial is a hard failure rather than a matter of tone.
The rule we hold ourselves to is simple: if a teammate could not make a change with their own login, the agent cannot make it for them either. And a denied action should stop — not retry, not switch environment to find a way around the block, not quietly report success.
5. Can you see the change before your customers do?
A screenshot in a chat window is not a preview. A preview is a URL you can click through, test a form on, open on your own phone and send to a colleague who has never logged into the tool. That difference decides whether approval is a real decision or a polite formality.
For anything that touches code or templates, add one more requirement: the assistant should verify the build actually succeeded before telling you the change is live. A commit that does not build never reaches the site, and being told otherwise is worse than being told nothing.
6. Is there a log you could hand to an auditor?
Attribution turns a scary change into a cheap one, because the worst case becomes a single click back to where you were. The log should record who asked, what changed field by field, when, in which environment, and who approved it — and it should be append-only, including for administrators. A log that privileged users can tidy up provides no assurance to anyone.
This is not only a compliance point. It is the reason a team stops being frightened of making changes at all.
7. Does it stop at content?
Plenty of real website jobs are not content jobs. Adding a section to a template needs code. Getting the right image needs sourcing, cropping and compressing. Answering which of our pages are quietly rotting needs the content graph joined to analytics. If the assistant can only edit text in fields, each of those still becomes someone else's ticket — and the queue you were trying to escape reassembles itself.
The unlock is having the thing that measures also be the thing that can edit. A report that ends in forty pages missing meta descriptions can then end in forty drafted meta descriptions, in the same conversation, for your approval.
A short evaluation script
If you are trialling something this quarter, these five requests will tell you more than any demo:
Create an entry of one of your custom types and check whether every required field came back populated.
Ask for a change across every page containing a particular phrase, and see whether you are shown the scope first.
Ask it to publish something using an account without publish rights, and watch what it does when refused.
Ask which pages are missing meta descriptions, then ask it to fix them — and see whether step two is possible.
Make a change, then ask what changed and who approved it. Time how long the answer takes.
If you want to see how we have answered these for specific platforms, there are write-ups for a Sanity CMS AI assistant and a Webflow AI assistant, plus the option of simply outsourcing website updates entirely.
The thing worth optimising for
Almost none of what matters here is about how clever the model is. Model quality determines how good the work is; the architecture determines how bad a mistake can be. For anything touching a production website, that is the right division of labour — optimise the intelligence for the average case and engineer the guardrails for the worst one.
Get it right and the day-to-day becomes pleasantly boring. Someone asks for a change, a draft appears, a preview gets reviewed, the right person approves, the site updates, the log records it. The assistant is not safe because it never errs. It is safe because the system around it makes errors small, visible and reversible — which is exactly why we trust human teams, and exactly the standard to hold a new one to.
See it work on your own site.
Connect your CMS, invite your team, and ship your first change on staging in two minutes.