# How to choose a GEO or AI visibility agency

> Nine questions that separate an agency with a method from one with a dashboard subscription and a content calendar. Ask them before you sign anything.

Source: https://armature.tech/library/choosing-an-ai-visibility-agency
Published: 2026-09-03
Publisher: Armature, Inc. (https://armature.tech)

---

The category is new enough that a lot of what is being sold is a dashboard subscription with a content calendar attached. Here is how to tell.

## The nine questions

### 1. What exactly will you measure, and how many prompts?

A serious answer names the metric and the sample. "Visibility across 200 tracked prompts on five engines, refreshed daily" is an answer. "AI visibility" is not.

Entry tiers on most tools cap at 50 prompts. A category with four use cases and four kinds of buyer needs more than that to be representative.

### 2. Will the prompt set include symptom wording?

This one separates people who have looked at real data from people who have not.

Users do not type category names. They type "people keep signing up with fake emails", not "recommend a bot protection service". Those two produce different answers and only one is usually in a tracked prompt set.

### 3. What will you change, not just report?

The whole game. A monthly report nobody acts on is a subscription.

Ask what the first three changes would be and who makes them. If the answer is "we will advise and you implement", price that honestly: you are buying a report and doing the work.

### 4. What happens if the number does not move?

Listen for a real answer. "We would look at which prompts moved and which did not, and at whether the pages we changed are being fetched at all" is a method. Silence is a dashboard.

### 5. How do you know the change caused the movement?

AI answers move for reasons nobody controls. A model updates and every number in your report changes.

An agency that attributes every rise to its own work and every fall to model drift is not measuring, it is narrating. Ask how they separate the two. A control group of untouched pages is the honest answer.

### 6. Do you measure the surface where my product is actually adopted?

If you sell a developer tool, this is the question that matters most and almost nobody will have an answer.

Prompt monitoring measures a person reading a chat answer. A large and growing share of developer tool adoption happens elsewhere: a coding agent inside a repository, given one line of instruction, choosing and installing something.

Those two come apart, measurably. The same hosted database request produced one winner in 111 of 111 JavaScript sessions and 48 of 132 TypeScript sessions. And on decision tasks, Claude Code ran a web search in 1.6% of runs against 53% for Codex, so published content is absent from most of one agent's decisions entirely.

An agency that has not thought about this is optimising half your market with full confidence.

### 7. What will you do about my documentation?

For a developer tool, the pages that decide adoption are the quickstart, the configuration reference, the limits page and the errors page. Not the blog.

An agency that wants to write thought leadership and does not want to touch your documentation is going to work on the wrong pages competently.

### 8. Who writes it, and do they understand the product?

Technical content written by somebody who has not run the code reads as such, to people and to models. Ask to see a sample in a comparable technical domain, and read the code in it.

### 9. What is the exit?

If you stop paying in six months, what do you keep? Pages you own, on your domain, are an asset. A dashboard login is not.

## What a good engagement looks like

| | Sign of a method | Sign of a subscription |
| --- | --- | --- |
| Measurement | Named metric, stated sample, symptom wording included | "AI visibility" |
| Action | They change things | They advise, you implement |
| Attribution | A control group, or an honest "we cannot fully separate this" | Every rise is their work |
| Scope | Documentation included | Blog only |
| Failure | A stated plan for the number not moving | Not discussed |
| Exit | Pages you own | A login |

## Red flags

**A promised ranking or share of voice number.** Nobody controls what a model says. And the score depends entirely on the prompt set: pick prompts your product is obviously right for and the number rises without anything improving.

**"We will get you into ChatGPT."** Nobody has that switch. What exists is being in the Bing index, being fetchable, and being the clearest answer.

**Anything about writing text addressed to the model.** It does not survive a person reading it, providers detect it, and it addresses none of the reasons products actually lose.

**Volume as the offer.** Forty posts that answer nothing specific lose to six pages that answer exactly what gets searched.

## What to do before you hire anyone

Run the cheap version yourself, because it changes which agency you need.

1. Ask the assistants your category's twenty real questions and record the answers. An afternoon.
2. Run 300 coding agent sessions: three repositories that look like your users' projects, ten requests written as symptoms, five runs each with Claude Code and Codex. A weekend and a few hundred dollars in tokens.
3. Read the losses.

If the losses are in the chat window, hire for generative engine optimization. If agents are installing your competitors while no human compares anything, no amount of that reaches the decision, and you need [a different measurement](/library/how-to-measure-install-share).

Knowing which of those you have is worth more than any agency's first three months.

## Common questions

### How do I choose a GEO agency?

Ask what they will measure, how many prompts, whether the prompt set includes symptom wording, what they will change rather than report, and what happens if the number does not move. An agency that cannot answer the last one is selling a dashboard.

### What should a GEO agency cost?

Anywhere from a few thousand a month to enterprise pricing. The price matters less than whether the engagement includes changing things. A retainer that produces a monthly report and no changes is a subscription to a chart.

### What is the biggest red flag?

Promising a ranking or a share of voice number. Nobody controls what a model says, the score depends entirely on the prompt set chosen, and a number that can be moved by changing the prompts is not a result.

### Does an agency need to understand developer tools specifically?

If you sell one, yes. The pages that decide adoption are the configuration reference and the quickstart, and the buyer is increasingly a coding agent rather than a person. A generalist brand agency will optimise the wrong pages competently.

## Read next

- [The best AI visibility tools for developer tools in 2026](https://armature.tech/library/best-ai-visibility-tools-for-developer-tools) (Markdown: https://armature.tech/library/best-ai-visibility-tools-for-developer-tools.md)
- [A developer marketing agency vs agent discoverability](https://armature.tech/library/developer-marketing-agency-vs-agent-discoverability) (Markdown: https://armature.tech/library/developer-marketing-agency-vs-agent-discoverability.md)
- [Is GEO worth it for a developer tool company?](https://armature.tech/library/is-geo-worth-it-for-a-developer-tool) (Markdown: https://armature.tech/library/is-geo-worth-it-for-a-developer-tool.md)
- [How to measure install share](https://armature.tech/library/how-to-measure-install-share) (Markdown: https://armature.tech/library/how-to-measure-install-share.md)

---

Armature helps software products get discovered and used by coding agents.
Service: https://armature.tech/discoverability · Results: https://armature.tech/leaderboards/sectors · Contact: contact@armature.tech
