🎩 You're Invited:Meet the Socket team at Black Hat in Las Vegas, August 3-6.RSVP
Sign In

dataraum

Package Overview
Dependencies
Maintainers
1
Versions
2
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

dataraum

Rich metadata context engine for AI-driven data analytics

pipPyPI
Version
0.2.1
Weekly downloads
18
-47.06%
Maintainers
1
Weekly downloads
 

DataRaum Context Engine

PyPI version Python License CI

A rich metadata context engine for AI-driven data analytics.

Traditional semantic layers tell BI tools "what things are called." DataRaum tells AI "what the data means, how it behaves, how it relates, and what you can compute from it."

The core insight: AI agents don't need tools to discover metadata at runtime. They need rich, pre-computed context delivered in a format optimized for LLM consumption.

Quick Start — MCP Server

The most common way to use DataRaum is as an MCP server inside Claude Desktop (or any MCP-compatible client).

# Install
pip install dataraum

# Or with uv
uv pip install dataraum

Add to your Claude Desktop config (claude_desktop_config.json):

{
  "mcpServers": {
    "dataraum": {
      "command": "dataraum-mcp"
    }
  }
}

Then in Claude Desktop:

Add the CSV files in /path/to/my/data and measure data quality

The server runs a 17-phase analysis pipeline and makes these tools available:

ToolDescription
begin_sessionStart an investigation session with a contract
add_sourceRegister a data source (CSV, Parquet, JSON, or directory)
lookExplore data structure, relationships, and semantic metadata
measureMeasure entropy scores, readiness, and data quality
queryNatural language query against the data
run_sqlExecute SQL directly with export support
end_sessionArchive workspace and end the session

Typical Workflow

add_source(name="accounting", path="/path/to/data")
  → begin_session(intent="explore data quality", contract="exploratory_analysis")
  → look()                    # Understand the data
  → measure()                 # Check quality scores and readiness
  → query("total revenue?")   # Ask questions
  → run_sql(sql="...", export_format="csv", export_name="report")
  → end_session(outcome="delivered")

Quick Start — CLI

# Run analysis pipeline (writes metadata.db + data.duckdb to ./pipeline_output)
dataraum run /path/to/data

# Inspect what was produced
dataraum dev context ./pipeline_output

See CLI Reference for all options.

What It Produces

DataRaum analyzes your data and generates:

  • Statistical metadata — distributions, cardinality, null rates, patterns
  • Semantic metadata — column roles, entity types, business terms (LLM-powered)
  • Topological metadata — relationships, join paths, hierarchies
  • Temporal metadata — granularity, gaps, seasonality, trends
  • Quality metadata — rules, scores, anomalies
  • Entropy scores — uncertainty quantification across all dimensions
  • Ontological context — domain-specific interpretation (financial, marketing, etc.)

LLM Configuration

Semantic analysis requires an Anthropic API key:

export ANTHROPIC_API_KEY="sk-..."

Configure the LLM provider in config/llm/config.yaml. See Configuration for details.

Development

git clone https://github.com/dataraum/dataraum
cd dataraum

# Install with dev dependencies (using uv)
uv sync --group dev

# Run tests
uv run pytest --testmon tests/unit -q

# Type check
uv run mypy src/

# Lint
uv run ruff check src/
uv run ruff format --check src/

Documentation

License

Apache 2.0 — see LICENSE.

Keywords

ai

FAQs

Did you know?

Socket

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

Related posts