Sign in
mcp server · aidatatools-dev.github.io

Data Quality Gate - deterministic post-scrape cleaner + verdict

Post-scrape data cleaner, no LLM: repairs mojibake, HTML, invisible chars. Plus a verdict.

What it says

The record the registry holds

Where it answers
https://www.aidatatools.dev/api/mcp_server · streamable-http
the record Copied from the official MCP registry, exactly as it holds it.
{
  "server": {
    "$schema": "https://static.modelcontextprotocol.io/schemas/2025-12-11/server.schema.json",
    "name": "io.github.aidatatools-dev/data-quality-gate",
    "description": "Post-scrape data cleaner, no LLM: repairs mojibake, HTML, invisible chars. Plus a verdict.",
    "title": "Data Quality Gate - deterministic post-scrape cleaner + verdict",
    "version": "0.3.0",
    "websiteUrl": "https://www.aidatatools.dev/llms.txt",
    "remotes": [
      {
        "type": "streamable-http",
        "url": "https://www.aidatatools.dev/api/mcp_server"
      }
    ],
    "_meta": {
      "io.modelcontextprotocol.registry/publisher-provided": {
        "boundary": "20 rules, published in full and free at GET https://www.aidatatools.dev/api/clean -- readable before paying. 7 are applied automatically (information-preserving), 5 need an explicit opt-in (they change row count, type or schema), and 8 are only ever reported, with no option to turn them on: near-duplicates are never merged, ambiguous placeholders are never nulled ('None' is a surname, 'NA' is Namibia), full NFKC is never applied (it rewrites 10^2 to 102), and failed extractions ('captcha', 'access denied') are never deleted -- that value is what tells you the record must be re-scraped. Impossible values are never repaired at all.",
        "docs": {
          "boundary": "https://www.aidatatools.dev/api/clean",
          "full": "https://www.aidatatools.dev/llms-full.txt",
          "openapi": "https://www.aidatatools.dev/openapi.json",
          "summary": "https://www.aidatatools.dev/llms.txt"
        },
        "guarantees": "Deterministic and idempotent. No value is emptied, retyped or reordered by an automatic rule. A repair that would introduce a defect the original did not have is reverted and flagged. The audited tier's ledger restores the input byte for byte.",
        "measured": "Real 100-row Amazon scrape, 57,956 nested string values: 453 repairs, 0 reverts, 0 values emptied. A control set of legitimate-but-suspicious data (French accents, Persian ZWNJ, emoji ZWJ, 10^2, AT&T, surname 'None', 'NA') comes back byte-for-byte identical: 0 changes, 11 flags.",
        "positioning": "Deterministic post-scrape data cleaner and quality gate. Repairs how data was ENCODED -- residual HTML, mojibake ('Café' -> 'Café'), invisible characters, non-breaking spaces -- and never what it SAYS. No LLM: identical input always yields byte-identical output, so results can be cached, replayed and audited.",
        "tags": [
          "data cleaning",
          "scraping repair",
          "post-scrape sanitization",
          "deterministic data repair",
          "mojibake correction",
          "encoding repair",
          "scraper output cleaning",
          "auditable data repair"
        ],
        "tools": {
          "check_dataset_quality": "FREE here. The full quality verdict: score, per-check facts, and RELIABLE / USABLE_WITH_CLEANING / UNRELIABLE. Bought once per source, then cached.",
          "clean_scraped_data": "Free to call, but it does NOT return the repaired data -- its response has no 'data' key. It reports which rules would fire on your data and what needs your decision, then names the paid call that hands the data back: POST https://www.aidatatools.dev/api/clean, $0.04 via x402, no account or signup. Call it after every extraction run -- a verdict is cached per source, but dirt is produced fresh by every run.",
          "clean_scraped_data_audited": "Same, pointing at POST https://www.aidatatools.dev/api/clean/audit, $0.12: the repaired data plus a replayable, reversible ledger of every transformation."
        },
        "why_repair_is_paid": "Detection is free on this endpoint and always has been. The repaired artifact is the paid product, because it is re-bought on every extraction run rather than cached like a verdict. Both cleaning tools say so in their own descriptions, before they are called."
      }
    }
  },
  "_meta": {
    "io.modelcontextprotocol.registry/official": {
      "status": "active",
      "statusChangedAt": "2026-08-12T07:23:32.533208Z",
      "publishedAt": "2026-08-12T07:23:32.533208Z",
      "updatedAt": "2026-08-12T07:23:32.533208Z",
      "isLatest": true
    }
  }
}
Make it installable

Draft an Agent Plugin from this

An MCP server is raw capability: some tools, wired to something. An Agent Plugin is the packaging that says what job it does, what it leaves behind and where it stops. The prompt below carries this record and asks for the packaging; your own assistant writes it, and nothing here is sent anywhere.

Make it installable

Draft an Agent Plugin from Data Quality Gate - deterministic post-scrape cleaner + verdict

Paste it into your assistant. It asks for the manifest, the server wiring and the skills, and for an honest account of what this record does not say. Read that second file first.

130 lines · the record is inside it, so nothing else is needed
Draft an Agent Plugin (agent-plugins.org, specification 1.1.0) that wraps the
MCP server described below, so that somebody could install one thing and have
an assistant that knows when and how to use it.

An Agent Plugin is one installable unit: a `plugin.json` manifest, an
`mcp.json` that wires up the servers it needs, and a `skills/` directory
where each skill is a folder holding a `SKILL.md`. Hand back every file in
full, each under its own path, ready to save.

1. Write `plugin.json` with `$schema` exactly `https://agent-plugins.org/schemas/1.1.0/plugin.schema.json`. The name is
   1 to 64 characters of a-z, 0-9, `-` and `.`, alphanumeric at both ends, with
   no `--` and no `..` in it.

2. Write `mcp.json` wiring THIS server exactly as its record declares it. A
   remote keeps the URL and the transport type as written. A package keeps the
   registry, the identifier and the version as written. Do not invent a command,
   a port, a flag or an argument that is not in the record.

3. Every secret stays an input. No key, token, password or connection string
   belongs in either file. Declare what has to be supplied, name it, and say what
   it is for.

4. Do not invent tools. The record lists the tools it lists, and if it lists
   none then the honest plugin says the tool list was not published rather than
   guessing one from the description.

5. Skills are jobs, not tools. Write one skill per thing somebody would actually
   ask for, and inside each one say when to reach for this server, what a good
   result looks like, and what to do when it comes back empty. A skill per tool
   is a manual page with a different filename.

6. Say where it stops. Name what this plugin will not do — what it has no tool
   for, what needs a person, and what it must not be pointed at. A plugin with no
   stated edge reads as one with no edge.

7. Keep the author's own words for the description. If you would rather say it
   differently, say yours somewhere else and leave theirs where it is.

8. Record which version of the server you wrapped, and where the record came
   from, at the top of `plugin.json`'s description or in the readme. A plugin
   nobody can trace back to a version is one nobody can update.

Produce a second file alongside them, `LIMITS.md`, and treat it as the more
important of the two. The plugin is for whoever installs it. This is for
whoever has to decide whether installing it is a good idea, and that is
usually a different person who will never read the manifest.

It has three parts.

**What this is built from.** One paragraph: whose server it is, what the
record says it does, which version, and the fact that the record is all you
had. Say plainly that nobody ran it.

**What the record does not say.** One entry per gap. Whether the tool list was
published. What the server does with what it reads. What it costs. Whether it
writes anything anywhere. What credentials it will ask for and what those
credentials can reach. An unanswered question stays an unanswered question:
do not fill one in from the description or from what similar servers usually do.

**What a person has to check before trusting it.** The specific things
somebody should verify for themselves, in the order that would stop them
soonest if the answer is bad.

Write it in plain English, and do not soften it. A plugin drafted from a
directory record is a starting point to argue with, not something to install
into anything that matters.

The server record follows, exactly as the public index holds it. It is
everything I have: nobody has run this server, called a tool on it, or checked
that the address answers.

```json
{
  "server": {
    "$schema": "https://static.modelcontextprotocol.io/schemas/2025-12-11/server.schema.json",
    "name": "io.github.aidatatools-dev/data-quality-gate",
    "description": "Post-scrape data cleaner, no LLM: repairs mojibake, HTML, invisible chars. Plus a verdict.",
    "title": "Data Quality Gate - deterministic post-scrape cleaner + verdict",
    "version": "0.3.0",
    "websiteUrl": "https://www.aidatatools.dev/llms.txt",
    "remotes": [
      {
        "type": "streamable-http",
        "url": "https://www.aidatatools.dev/api/mcp_server"
      }
    ],
    "_meta": {
      "io.modelcontextprotocol.registry/publisher-provided": {
        "boundary": "20 rules, published in full and free at GET https://www.aidatatools.dev/api/clean -- readable before paying. 7 are applied automatically (information-preserving), 5 need an explicit opt-in (they change row count, type or schema), and 8 are only ever reported, with no option to turn them on: near-duplicates are never merged, ambiguous placeholders are never nulled ('None' is a surname, 'NA' is Namibia), full NFKC is never applied (it rewrites 10^2 to 102), and failed extractions ('captcha', 'access denied') are never deleted -- that value is what tells you the record must be re-scraped. Impossible values are never repaired at all.",
        "docs": {
          "boundary": "https://www.aidatatools.dev/api/clean",
          "full": "https://www.aidatatools.dev/llms-full.txt",
          "openapi": "https://www.aidatatools.dev/openapi.json",
          "summary": "https://www.aidatatools.dev/llms.txt"
        },
        "guarantees": "Deterministic and idempotent. No value is emptied, retyped or reordered by an automatic rule. A repair that would introduce a defect the original did not have is reverted and flagged. The audited tier's ledger restores the input byte for byte.",
        "measured": "Real 100-row Amazon scrape, 57,956 nested string values: 453 repairs, 0 reverts, 0 values emptied. A control set of legitimate-but-suspicious data (French accents, Persian ZWNJ, emoji ZWJ, 10^2, AT&T, surname 'None', 'NA') comes back byte-for-byte identical: 0 changes, 11 flags.",
        "positioning": "Deterministic post-scrape data cleaner and quality gate. Repairs how data was ENCODED -- residual HTML, mojibake ('Café' -> 'Café'), invisible characters, non-breaking spaces -- and never what it SAYS. No LLM: identical input always yields byte-identical output, so results can be cached, replayed and audited.",
        "tags": [
          "data cleaning",
          "scraping repair",
          "post-scrape sanitization",
          "deterministic data repair",
          "mojibake correction",
          "encoding repair",
          "scraper output cleaning",
          "auditable data repair"
        ],
        "tools": {
          "check_dataset_quality": "FREE here. The full quality verdict: score, per-check facts, and RELIABLE / USABLE_WITH_CLEANING / UNRELIABLE. Bought once per source, then cached.",
          "clean_scraped_data": "Free to call, but it does NOT return the repaired data -- its response has no 'data' key. It reports which rules would fire on your data and what needs your decision, then names the paid call that hands the data back: POST https://www.aidatatools.dev/api/clean, $0.04 via x402, no account or signup. Call it after every extraction run -- a verdict is cached per source, but dirt is produced fresh by every run.",
          "clean_scraped_data_audited": "Same, pointing at POST https://www.aidatatools.dev/api/clean/audit, $0.12: the repaired data plus a replayable, reversible ledger of every transformation."
        },
        "why_repair_is_paid": "Detection is free on this endpoint and always has been. The repaired artifact is the paid product, because it is re-bought on every extraction run rather than cached like a verdict. Both cleaning tools say so in their own descriptions, before they are called."
      }
    }
  },
  "_meta": {
    "io.modelcontextprotocol.registry/official": {
      "status": "active",
      "statusChangedAt": "2026-08-12T07:23:32.533208Z",
      "publishedAt": "2026-08-12T07:23:32.533208Z",
      "updatedAt": "2026-08-12T07:23:32.533208Z",
      "isLatest": true
    }
  }
}
```

The server is listed as "Data Quality Gate - deterministic post-scrape cleaner + verdict".
What this is

MCP server

application/mcp-server-record+jsonhosted · an address you callversion 0.3.0last seen 2026-08-23

These are the publisher's own words, filed by them with the official MCP registry and mirrored here. This catalog did not read them from the publisher and has not run, called or installed anything.

Where it lives

The publisher's own address

https://www.aidatatools.dev/llms.txt

The address the publisher put in their registry record, printed as they wrote it. Nobody here has opened it.