
Company News
AWS Security Hub Adds Socket for Supply Chain Security
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.
data-profiler-mcp
Advanced tools
MCP server that profiles tabular data files (CSV, TSV, Parquet, Excel, JSON): schema, distributions, data-quality flags, and dtype suggestions for LLM agents.
An MCP server that lets an LLM understand any tabular data file: point it at a CSV, Parquet, Excel or JSON file and get schema, distributions, data-quality flags and dtype suggestions back as structured JSON.
Stop pasting df.head() and df.info() into chat. Ask your assistant "profile sales.csv" and it reads the file itself, then tells you what is in it, what is wrong with it, and how to load it more efficiently.

Works with Claude Desktop, Claude Code, Cursor, or any MCP-compatible client.
Six focused tools, all returning clean JSON:
| Tool | What it does |
|---|---|
profile_dataset | One-call overview: shape, memory, missing-value summary, duplicate rows, a per-column summary, and plain-language quality flags. |
preview_data | The first / last / a random sample of n rows as real records. |
column_stats | Deep dive on one column: full percentiles, skew/kurtosis, outliers (IQR), a histogram, or top values + string lengths for text. |
detect_quality_issues | A data-quality audit: duplicates, high-missing and constant columns, numbers stored as text, mixed-type columns, whitespace padding, likely IDs, grouped by severity. |
suggest_dtypes | Memory-saving / type-fixing recommendations (text to numeric, low-cardinality to category, integer/float downcasting) with estimated savings. |
compare_datasets | Diff two files: added/removed columns, dtype changes, row-count delta, and per-column null-rate and mean side by side. |
Supported formats: CSV, TSV, Parquet, Excel (.xlsx/.xls), JSON and JSON Lines. Large files are read up to a row cap and clearly flagged as sampled.
No dataset at hand? examples/sample.csv is a small sales export with deliberate quality issues (missing regions, a duplicate row, a constant column, whitespace padding) -- ask your assistant to "profile examples/sample.csv" and see what it flags.
Requires Python 3.10+.
# with uv (recommended)
uv tool install data-profiler-mcp
# or with pip
pip install data-profiler-mcp
Or run it straight from source without installing:
git clone https://github.com/haiiibin/data-profiler-mcp
cd data-profiler-mcp
uv run data-profiler-mcp
Edit claude_desktop_config.json
(macOS: ~/Library/Application Support/Claude/, Windows: %APPDATA%\Claude\) and add:
{
"mcpServers": {
"data-profiler": {
"command": "data-profiler-mcp"
}
}
}
Running from source instead of installing? Point it at the checkout:
{
"mcpServers": {
"data-profiler": {
"command": "uv",
"args": ["--directory", "/absolute/path/to/data-profiler-mcp", "run", "data-profiler-mcp"]
}
}
}
Restart Claude Desktop and the tools appear under the plug icon.
claude mcp add data-profiler -- data-profiler-mcp
Once connected, just talk to your assistant:
~/data/sales_2025.csv and tell me what's in it."customers.parquet?"events.jsonl."revenue column, including outliers."snapshot_jan.csv and snapshot_feb.csv?"profile_dataset{
"file": { "name": "sample.csv", "format": "csv", "size_human": "14.2 KB" },
"shape": { "rows": 201, "columns": 13, "sampled": false },
"memory_usage_human": "78.4 KB",
"missing_summary": { "total_missing_cells": 561, "pct_missing": 21.5, "columns_with_missing": 3 },
"duplicate_rows": { "count": 1, "pct": 0.5 },
"columns": [
{
"name": "price", "dtype": "float64", "inferred_type": "float",
"non_null": 201, "null": 0, "unique": 51,
"stats": { "min": 0.0, "max": 100000.0, "mean": 521.3, "median": 24.0 }
}
],
"quality_flags": [
"[high] empty_col: Column is entirely empty (all values missing).",
"[warning] const: Column holds a single constant value; it carries no information.",
"[warning] numeric_text: Every value parses as a number but the column is stored as text."
]
}
detect_quality_issues{
"issue_count": 8,
"severity_counts": { "high": 2, "warning": 4, "info": 2 },
"issues": [
{ "column": "empty_col", "issue": "all_missing", "severity": "high",
"detail": "Column is entirely empty (all values missing)." },
{ "column": "numeric_text", "issue": "numeric_stored_as_text", "severity": "warning",
"detail": "Every value parses as a number but the column is stored as text." }
]
}
The server is built on FastMCP and reads files with pandas (plus pyarrow for Parquet and openpyxl for Excel). Every tool returns a plain, JSON-serializable dict, with NumPy scalars, NaN/inf and timestamps normalized so the output is safe to hand straight back to a model. Nothing is written to disk and no data leaves your machine.
uv venv
uv pip install -e ".[dev]"
uv run pytest
MIT. See LICENSE.
FAQs
MCP server that profiles tabular data files (CSV, TSV, Parquet, Excel, JSON): schema, distributions, data-quality flags, and dtype suggestions for LLM agents.
We found that data-profiler-mcp demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Company News
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.

Research
/Security News
Popular npm packages keyv and cacheable compromised.

Security News
A misconfiguration gave three Anthropic models internet access, and one, believing it was in a simulation, shipped a credential-stealing package to PyPI.