mudpie

practical-guide · 5 min read

How to measure agent traffic without counting the wrong thing

A worksheet for separating crawler requests, page reads, tool use, and verified outcomes.

Published · Updated

Count the job the agent finished.

A crawler can fetch a thousand pages without recommending your product. An agent can read one page, answer its user, and leave. Both can appear in a traffic report. They call for different decisions.

I would start with a worksheet for one task: finding a price, comparing two plans, or completing a booking. Keep the task narrow enough that you can check the result yourself.

What each event earns you

Recorded event What you can say What remains unknown
Crawler request A client requested a URL Whether a person asked for it or an answer used it
Request classified as an agent read A request met your recorded classification rules Whether the identity is genuine or the content was used
Tool registration A page made a tool available Whether an agent discovered or invoked it
Tool invocation A client attempted an operation Whether the result was correct or useful
Successful response The operation returned the expected success state Whether the user's task was completed
Verified outcome An inspected answer, page state, or destination record satisfies the task Whether the agent caused an incremental sale

These are proposed reporting definitions. Write down your own classifier and success rules beside the chart. A label such as “agent visit” hides too much if nobody can explain what qualifies.

Copy this worksheet

Field Fill this in before the change
Task What should the user be able to accomplish?
Eligible operations Which task and population qualify before you inspect success or failure?
Success evidence What record, page state, or receipt proves completion?
Failure evidence What would show a wrong answer, abandoned task, or rejected action?
Exclusions Which owner tests, scanners, retries, and probes are excluded?
Comparison Which complete dates and unchanged routes will you compare?
Missing data Which routes or clients are invisible to your collector?
Decision What result would make you keep, revise, or remove the change?

For a booking task, a tool response saying “done” is weak evidence. Check that the booking exists with the requested date and service. For a research task, check whether the returned source supports the answer. Do not borrow the booking definition for research because it produces a cleaner conversion rate.

Use denominators that match the question

For a verified-completion rate, divide completed logical operations by eligible logical operations. Keep failed and unknown outcomes in the denominator. Report the failure and unknown shares beside the result. This measures what you could verify; missing outcome evidence can make it lower than actual completion. Count invocation attempts separately when assessing retry load. Deduplicate retries only when you have a reliable operation identifier; matching text alone can merge different users asking the same question.

For content discovery, count first entries to the page separately from reads later in a journey. A new guide can collect reads from your old article without bringing another visitor to the site. Inspect both URLs and the topic total before claiming growth.

For adoption, separate your own tests from unprompted use. A successful test shows that the route works under those conditions. It does not establish that someone chose to use it.

Work one small sample by hand

Here is a fictional fixture, not customer data. There are 15 invocation records: three tagged owner tests, two retries of an already counted operation, and ten distinct eligible operations. Within the fixed observation window, six have verified results, two have known failures, and two have no outcome evidence. Report 6/10 verified completions, 2/10 failures, and 2/10 unknowns. The 12 non-test invocations are a separate attempt count. You cannot call this an 80% success rate by quietly dropping the unknowns.

For passive requests, keep identity evidence beside the count. A request with an agent-like user-agent header can be labeled declared agent, unverified. A tagged owner request belongs in tests even if the header looks genuine. An unrecognized request stays unknown. Do not upgrade either to verified identity because the page was fetched successfully. Your traffic report should expose these classes rather than hide them in one total.

I would rather leave the decision open than pick a winner from two requests. When a few operations determine the result, show their counts and failures, extend the collection window, and keep the same eligibility rules.

Three ways to fool yourself

A partial day looks like a decline. Compare complete days in one timezone. Record outages, launches, and distribution changes beside the comparison.

A better collector looks like growth. If a release starts observing previously invisible requests, preserve that boundary. Report the old and new coverage rather than drawing one continuous trend through it.

An unknown outcome disappears. Keep an unknown bucket. If you cannot connect a handoff to signup, report the handoff and the missing link. Counting it as a conversion changes the question after the fact.

Chrome describes WebMCP as a way for pages to expose structured tools. Registration and task completion are separate steps in that design. Its MCP comparison also distinguishes live-page tools from server-side access. Record those access paths separately before combining any totals.

Use Mudpie to inspect the activity it records, then fill in the worksheet with your product's outcome evidence. Leave a cell marked unknown when the evidence is missing.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗