mudpie

practical-guide · 5 min read

A WebMCP analytics implementation checklist

Verify tool discovery, invocation, failures, and outcomes without counting your own tests as adoption.

Published · Updated

Start with a task that is failing. A missing answer may need a better page; a blocked account may need a permissions repair. When the task needs a tool, test one operation from discovery through its result before adding more.

Chrome's WebMCP documentation describes page-defined tools and distinguishes registration from execution. Browser and client support can change; check the current documentation and record the versions used in your test.

This checklist is an implementation and measurement proposal. Event names below are examples, not a Mudpie API contract.

Use the checklist with your agent

Get the WebMCP engineering skill on GitHub. It covers tool design, state changes, human handoff, and outcome checks, drawing on Chrome, MCP-B, Stripe, Sierra, and Decagon. Mudpie is optional.

npx skills add aboul3ata/mudpie-agent-kit --skill webmcp-engineering

Ask your agent: “Use webmcp-engineering to add a useful browser tool to this app. Reuse the existing product logic and verify the result.”

Define the operation

Pick a task whose outcome you can inspect: find a plan's limits, filter a catalog, or create a draft. Write down the inputs, allowed account scope, expected result, and failure behavior.

Start with a read-only operation when it answers a real user need. For actions that change state, retain the product's authorization and confirmation requirements. A tool interface should reach the same permission checks as the UI.

Keep four records separate

Suggested record Minimum useful fields Verification
Tool available Page, tool name, schema version, timestamp Registration was recorded; client discovery requires its own test
Invocation started Operation ID, attempt ID, tool, access path, timestamp The handler received the request
Invocation finished Operation ID, attempt ID, result class, duration The result matches success or failure rules
Outcome verified Operation ID, outcome type, verification time The destination contains the expected result

The caller creates one logical operation ID for a user-authorized action and a new attempt ID for each invocation. Propagate both through the handler and collector; attach the operation ID to the destination receipt when supported. A retry keeps the operation ID and gets a new attempt ID.

Use a result class such as success, invalid input, unauthorized, cancelled, timeout, or server error. Preserve the difference between a failure and an outcome you could not verify.

Record the page route without credentials or private query strings. Collect only the request context needed for the analysis. Full prompts can contain private information; logging them by default creates a second data-retention problem.

Write down one trace

This is an illustrative trace for a read-only catalog search, not a captured production run. The caller supplies operation op-17. The handler records attempt a-1 when execution begins, records a timeout when the response deadline expires, then records attempt a-2 for the retry. The second attempt returns three products. A separate check confirms those products satisfy the requested filters and records the verified result against op-17. That is two attempts and one completed operation.

A safe start event could contain operation_id, attempt_id, tool: catalog_search, route: /catalog, access_path: webmcp, and a timestamp. The finish event adds a result class and duration. Omit cookies, authorization headers, full page URLs with private queries, and the user's full prompt. The handler emits execution events; the outcome verifier emits its own result after inspection. Send telemetry through a bounded, non-blocking collector so its failure cannot prevent the product operation.

Use the current Chrome setup and API documentation to wire the tool. The event design here is independent of the browser registration syntax and does not imply Mudpie automatically stores these custom fields.

Run these checks

  1. Open the page in the target client. Confirm that the intended tool is discoverable. Repeat after navigation or a workspace change if the tool depends on that state.
  2. Invoke it with valid inputs. Inspect both the returned answer and the destination result.
  3. Try an invalid input and a disallowed account scope. Confirm a useful failure and no unauthorized result.
  4. Interrupt or cancel an operation. Check that the UI, returned status, and actual state agree. A timeout does not establish that a write failed.
  5. Retry using the same operation identifier where the operation supports idempotency. Confirm that a write is not duplicated.
  6. Check collection separately. Match the test's operation ID across the start, finish, and outcome records. An HTTP acknowledgment from a collector alone does not prove durable storage.
  7. Tag the run as a test, retain the evidence privately, and remove the disposable records it created. Verify cleanup.

Compare access paths without double counting

A browser tool and a server-side MCP tool can call the same backend operation. Keep the access path on the event. Count the operation once only when both paths share a reliable logical operation ID. Otherwise report separate totals and unknown overlap; matching inputs do not prove it is the same action. Do not add a page load, invocation, and booking together and call the sum customers.

Chrome's comparison guide explains the live-page boundary. Use that distinction when deciding which state belongs in the browser and which permissions and business rules belong in the service.

Put an explicit limit on the result

A useful release note says which client and version completed which task, under which account conditions, and what failed. “WebMCP works” leaves the next person guessing.

After release, inspect unprompted attempts separately from this test run. Use the measurement worksheet to define success and the content decision guide when failures reveal missing public information. Open Mudpie to inspect the activity recorded for your site.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗