The endpoint is live, and so is signing in. https://fastopendata.com/mcp is deployed, and the authorization server behind it issues real tokens against your FastOpenData account — the device grant, so no browser redirect and no callback port is involved. Two caveats, both specific. The fastopendata-mcp bridge is not on PyPI yet, so the uvx lines below need the repository until it is; you can do the login with curl in the meantime, and §2 shows how. And a hosted client that expects to redirect you through a browser — a remote connector in somebody else's product — cannot complete a login yet: we advertise no authorization endpoint for it to redirect to, and it will say so at discovery rather than failing halfway. Everything else below is the server as implemented. Dates on this page would be guesses, so there are none.
Model Context Protocol

FastOpenData for agents

One endpoint, seven tools, federal open data for every U.S. Census tract, county, PUMA and state. An agent asks for variables by name and gets the rows it asked for — not a 900-column dump to reason through. The same data is available to programs over the REST API; the two interfaces meter differently — requests there, cells here.

MCP 2026-07-28 · Streamable HTTP, POST only · OAuth 2.1, audience-bound

1. The setup block

Two ways in. For a host that speaks remote MCP, the whole configuration is a URL and a token:

{
  "mcpServers": {
    "fastopendata": {
      "url": "https://fastopendata.com/mcp",
      "headers": { "Authorization": "Bearer ${FASTOPENDATA_TOKEN}" }
    }
  }
}

For a host that only speaks stdio, a thin bridge forwards to the same endpoint. It is written and tested but not published to PyPI yet, so uvx will not find it today — install it from the repo, or read this as the shape it will take:

{
  "mcpServers": {
    "fastopendata": {
      "command": "uvx",
      "args": ["fastopendata-mcp"],
      "env": { "FASTOPENDATA_TOKEN": "${FASTOPENDATA_TOKEN}" }
    }
  }
}
The token never goes in args. A command line is visible in ps, in shell history, and in whatever log your host keeps of how it started its subprocesses. fastopendata-mcp refuses to start if it finds something token-shaped in argv, rather than warning — a warning during an agent launch is a warning nobody reads.

With no FASTOPENDATA_TOKEN set, the bridge reads a stored credential instead. The login below is implemented in the bridge and waits on the same missing piece as everything else — there is no issuer to log in to yet:

uvx fastopendata-mcp login
# prints a URL and a short code to type into it — the OAuth 2.1 device
# grant, so it works over SSH and inside a container

uvx fastopendata-mcp status
# where the credential came from, which scopes, when it expires

The login writes ~/.config/fastopendata/token.json at mode 600 and refreshes the access token in the background while the bridge runs. FASTOPENDATA_TOKEN takes precedence over a stored login, so a CI job on your laptop uses the token it was given rather than your account.

2. Getting a token

The endpoint publishes where tokens come from, so a client that implements OAuth discovery needs nothing configured:

curl https://fastopendata.com/.well-known/oauth-protected-resource/mcp

# {"resource": "https://fastopendata.com/mcp",
#  "authorization_servers": ["https://fastopendata.com"],
#  "scopes_supported": ["fod:read", "fod:aggregate"],
#  "bearer_methods_supported": ["header"]}

We are our own identity provider: the issuer is https://fastopendata.com, its metadata is at /.well-known/oauth-authorization-server, and you sign in with the same username and password as the REST API. There is no separate identity to create, and no third party sees the password — the form that takes it is served from this domain. Read the issuer out of the resource document rather than hard-coding it; it being co-hosted today is an implementation detail.

Two grants, and one notable absence. The device grant (RFC 8628) and refresh rotation are supported; there is no authorization endpoint, which is what a hosted remote connector would need. Doing the device flow by hand is four calls, and it is exactly what the bridge does:

# 1. Register once. Public client, no secret, no authentication.
curl -sX POST https://fastopendata.com/oauth/register \
  -H 'content-type: application/json' \
  -d '{"client_name": "my script"}'        # -> client_id

# 2. Ask for a code pair.
curl -sX POST https://fastopendata.com/oauth/device_authorization \
  -d client_id=fod-… \
  -d 'scope=fod:read fod:aggregate' \
  -d resource=https://fastopendata.com/mcp

# 3. Open the verification_uri_complete it returns, read what you are
#    approving, and sign in. Then poll:
curl -sX POST https://fastopendata.com/oauth/token \
  -d grant_type=urn:ietf:params:oauth:grant-type:device_code \
  -d device_code=… -d client_id=fod-…

Before you approve, step 3 answers authorization_pending; poll faster than every five seconds and it answers slow_down; a code lives ten minutes. You get an access token good for an hour and a refresh token good for thirty days that rotates on every use. If a spent refresh token is ever presented again, every token from that login is revoked and you log in again — deliberately, because a token presented twice may have been copied and there is no way to tell which presenter is the real one.

If a code arrives that you did not ask for, do not type it in. The device flow is approved by reading a code off one screen and typing it into another, which is also how it would be abused: whoever started the login gets a token that reads data as you and spends your budget. Nobody from FastOpenData will ever send you a code. The approval page says so, and its "Don't approve" button needs no password, so stopping a login you did not start is one click.

Two scopes, and they are not the same purchase. fod:read covers every lookup tool — one place, or a handful at a time. fod:aggregate covers aggregate_variable, which reads many rows to return a few.

A token must be minted for this resource. Every token request carries resource=https://fastopendata.com/mcp (RFC 8707), and a token whose audience is something else is rejected here. The failure is confusing from outside: a token that is perfectly valid at the identity provider produces nothing but 401s, and the endpoint is working correctly while it does.

3. Your first call

There is no initialize handshake — ask the server what it is, then list the tools, then call one. One JSON-RPC request per POST: batches are refused, there is no SSE stream, and GET returns 405 with a body saying why.

curl -s https://fastopendata.com/mcp \
  -H "Authorization: Bearer $FASTOPENDATA_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"jsonrpc":"2.0","id":1,"method":"server/discover"}'
Coverage, before anything else: the address geocoder holds an OpenStreetMap extract for Georgia only. The tract, county, PUMA and state tables are national — it is the address lookup in front of them that is not. An address elsewhere fails at the geocode stage, and that is a coverage limit rather than missing data. describe_coverage says so, takes no arguments, and costs one cell.

4. The tools

Seven, of which five are listed today. tools/list is sorted by name and varies only with the scopes of the token you present.

ToolScopeTakesReturns
describe_coveragefod:readnothinggeographies, vintages, what the geocoder covers
search_variablesfod:readquery, limit, geo_levelranked variable names, best first
describe_variablefod:readnameone variable in full: units, range, source, vintage
lookup_locationfod:readan address, a point, or a geoidthe geoids containing that place
get_location_datafod:readgeoid, variablesone row, the named columns only
compare_locationsfod:readgeoids, variablesthe same columns across 2–25 places
aggregate_variablefod:aggregatevariable, group_by, within, filter, statisticsone row per group, up to 500 groups

compare_locations and aggregate_variable are withheld until the variable dictionary is complete. Shipping them against undocumented column names would teach the first users of this server the wrong interface.

Variables must be named. There is no "give me everything" argument on any tool, and get_location_data takes at most 40 names. This is the central design decision and not an oversight: a tool that returns 885 columns fills an agent's context with numbers it did not ask for, and the agent then reasons worse, not better. Start from search_variables. Names are prefixed by the level they are measured at — cre_ and nri_ are tract, then puma_, county_, state_.

5. Four worked templates

Each is a sequence of calls, written the way an agent should be told to run it.

Address enrichment

"What kind of neighbourhood is this address in?"

lookup_location   {"address": "1 W Court Sq, Decatur, GA"}
#  -> geoids: {tract: "13089020100", county: "13089",
#              puma: "1301504", state: "13"}, match_precision: 30

search_variables  {"query": "social vulnerability population",
                   "geo_level": "tract"}

get_location_data {"geoid": "13089020100",
                   "variables": ["cre_population",
                                 "nri_social_vulnerability",
                                 "puma_pct_bachelors_5yr"]}

Three calls, eight cells. A tract geoid answers coarser columns too: puma_pct_bachelors_5yr is a PUMA variable, and asking for it with a tract geoid returns the value for the PUMA containing that tract — there is exactly one. The reverse does not hold; a county geoid cannot answer a tract column, and asking gets a mismatch error rather than an average. Do not skip the search: a guessed column name is an error, not a near miss.

County comparison

"How do these five counties differ?"

search_variables  {"query": "rural urban population",
                   "geo_level": "county"}

compare_locations {"geoids": ["13121", "13089", "13067",
                              "13135", "13063"],
                   "variables": ["county_population_2020",
                                 "county_rucc_2023",
                                 "county_pres_two_party_votes_2024"]}

One call returns the matrix: 17 cells, two of overhead and one per value. Five separate get_location_data calls return the same numbers for 25 cells, spend five of your twenty calls a minute, and leave the agent to line the results up itself.

Risk screening

"Which tracts in this county look exposed?"

aggregate_variable {"variable": "nri_risk_score",
                    "group_by": "tract",
                    "within": "13121",
                    "filter": "cre_population > 1000",
                    "statistics": ["count", "mean", "p90"]}

get_location_data  {"geoid": "<the tract that stood out>",
                    "variables": ["nri_flood_risk_score",
                                  "cre_pct_high_vulnerability"]}

Needs fod:aggregate, and costs 25 cells flat however broad the question is — an aggregate is deliberately the cheap way to ask something wide, because charging per underlying row would push callers toward reading rows one at a time. Two rules make it work: within always, and the grouping no finer than the question. A result holds at most 500 groups, so a tract grouping with no within is refused — 86,200 groups — and the refusal says to group coarser or add a scope. A group backed by fewer than five rows is suppressed rather than shown. A query running past 15 seconds is stopped, with the same advice to narrow it.

Variable discovery

"What do you even have about risk?"

describe_coverage  {}

search_variables   {"query": "flood wildfire expected annual loss",
                    "limit": 25}

describe_variable  {"name": "nri_flood_risk_score"}

Worth doing once per project and keeping the names. describe_variable answers "is this a percentage or a proportion, and from which vintage" — the question that silently ruins analyses — and describe_coverage reports how many of the 885 columns are documented, so "that is not written down yet" is visible rather than inferred.

6. Budgets

Metered in cells: one variable at one geography is one cell. Plus a calls-per-minute ceiling, and a cap on how many distinct geographies one subject may read per rolling window — that last one is what makes copying the dataset through this interface take years rather than days.

TierCells / dayCalls / minDistinct tracts
free2,0002050
standard50,000120250
partnernegotiatednegotiatednegotiated

Fixed prices: describe_coverage, search_variables and describe_variable cost one cell; lookup_location costs two, because it may run the geocoder; a row read costs two plus one per value returned; an aggregate costs 25. A failed call still costs the overhead — free failures would make this a free oracle for "does this geoid exist".

These are the numbers the code enforces, not a price list: nothing is billed today, and what an account costs once it is billed is sketched on the pricing page and labeled as planned. Cells and requests are different meters, and an account carries both — the API counts requests, the agent interface counts cells.

A refused call names which budget and when it resets. It is a tool error, not a transport error: the agent sees it and should stop rather than retry. Retrying a budget refusal spends the calls budget on learning the same thing twice.

7. Deprecation policy

The tool list is a published contract. An agent caches it — most hosts put it straight into a prompt — and unlike a browser client it will not read a migration note. So:

Additive changes only, with a window.

  • A new tool, a new optional argument, or a new outputSchema field may appear at any time. Clients must tolerate unknown fields in results.
  • A tool name is never reused for different behaviour, and never silently removed.
  • Removing a tool, removing or renaming an outputSchema field, making an optional argument required, or narrowing an argument's accepted values is a breaking change and gets at least 180 days between announcement and removal.
  • During that window the tool keeps working, and its tools/list entry carries _meta["io.fastopendata/deprecation"] with since, sunset and, where there is one, replacement. The announcement travels with the thing being announced; there is no mailing list to miss.
  • Variable names are governed by the same rule. A variable is renamed only by adding the new name and deprecating the old one.
  • Bug fixes to values are not deprecations. A number that was wrong gets corrected, and the vintage in describe_variable says which snapshot it came from.

8. Caching the tool list

capabilities.tools.listChanged is false, and that is a fact about the transport rather than a preference: the endpoint is POST-only with no session and no stream, so there is no channel on which to push a notification. Claiming otherwise would promise a message that could never arrive.

What you get instead is a fingerprint. Every tools/list result carries _meta["io.fastopendata/tool-list-digest"] over exactly what was served, and _meta["io.fastopendata/deprecation-policy"] pointing back at this page. Store the digest; compare it; re-read the list when it moves. It changes when anything a client might have cached changes — a renamed field inside an outputSchema counts — and the order of the list never changes, because it is sorted by name.

9. When it goes wrong

Protocol faults are JSON-RPC error codes. Everything else is a tool error with isError: true and a message written to be read by the agent, because an agent can act on a sentence and cannot act on a code.

What you see
What it means
Do
"Narrow it and call again"
The query was stopped at its time bound.
Add within / filter
"out of … resets at …"
A budget, named, with a reset time.
Wait
geocode stage, Georgia only
Coverage, not an outage.
describe_coverage
401 on every call
Usually the audience: a token minted for another resource.
Re-mint
403 insufficient_scope
Valid token, missing scope — the scope is named.
Re-mint
503, geocoder
Address lookups are down; geoid reads still work.
Retry later
503 temporarily_unavailable
The deployment has no issuer, or the issuer has no signing key — failing closed, not refusing you.
Retry later

Honestly about capacity: this endpoint runs a single Cloud Run instance by deliberate choice, and the geocoder is one small VM. Agent traffic is burstier than browser traffic — one question can fan out to a dozen calls in a second — so a client should bound its own concurrency. The stdio bridge keeps at most four requests in flight for exactly that reason.

The same material, in more detail and with the operational side attached, lives in the repository: docs/mcp_connecting.md and docs/mcp_runbook.md.