Fast, minimal profiler for CSV and Parquet files.
DatasetPeek is a small server-rendered app built with Robyn, Polars, and Jinja2. It gives a technical user a quick first-pass read on a local, S3, or MinIO dataset without turning the UI into a full EDA tool.
- Local and S3-compatible CSV/Parquet profiling.
- Deterministic orientation summary and lightweight next-check guidance.
- Column role hints, quality signals, capped top values, and numeric summaries.
- Random sample cycling plus head and tail previews.
- Markdown and standalone HTML report downloads generated from the in-memory profile.
- Demo-friendly social preview metadata for shared app links.
Requirements:
- Python 3.12+
uv
Install dependencies:
make installStart the app:
make runOpen http://127.0.0.1:8080.
The launcher also respects PORT and HOST, which is useful for managed platforms:
PORT=9090 HOST=0.0.0.0 uv run python main.pyDatasetPeek accepts local CSV/Parquet uploads and S3-compatible object URIs. Use the source switcher on the home page to choose between a local file and an s3:// object:
s3://bucket/path/data.csv
For private AWS S3, MinIO, Cloudflare R2, or another S3-compatible object store, configure read-only credentials through environment variables:
DATASETPEEK_S3_ENDPOINT_URL=http://localhost:9000 # MinIO/custom S3/R2 endpoint
DATASETPEEK_S3_ACCESS_KEY_ID=minioadmin
DATASETPEEK_S3_SECRET_ACCESS_KEY=minioadmin
DATASETPEEK_S3_REGION=us-east-1
DATASETPEEK_S3_FORCE_PATH_STYLE=trueIf DATASETPEEK_S3_ENDPOINT_URL is set, DatasetPeek uses path-style requests such as http://localhost:9000/bucket/path/data.csv, which matches MinIO's default setup and many S3-compatible providers. Without credentials, DatasetPeek attempts anonymous reads, but public S3 bucket behavior is provider- and policy-dependent.
DatasetPeek profiles the full uploaded file or S3 object when it is within the size limit. If the object is an exported sample from a larger dataset, the reported rows, signals, and summaries describe that sample.
Legacy DATAPEEK_* environment variables are still accepted as fallbacks during the rename.
DatasetPeek reads operational settings from environment variables at runtime:
| Setting | Default | Purpose |
|---|---|---|
DATASETPEEK_MAX_UPLOAD_MB |
100 |
Reject uploads or S3 objects above this size. |
DATASETPEEK_LARGE_FILE_WARNING_MB |
50 |
Show a large-file warning above this size. |
DATASETPEEK_RANDOM_SAMPLE_ROWS |
10 |
Number of rows in each random sample preview. |
DATASETPEEK_HEAD_TAIL_ROWS |
5 |
Number of rows shown in head and tail previews. |
DATASETPEEK_SAMPLE_VALUE_COUNT |
3 |
Number of sample values shown per column. |
DATASETPEEK_TEXT_TRUNCATE_CHARS |
50 |
Maximum displayed length for cell/sample text. |
DATASETPEEK_TOP_VALUES_LIMIT |
5 |
Maximum top values shown for compact categorical/flag fields. |
DATASETPEEK_CSV_INFER_SCHEMA_ROWS |
5000 |
Number of CSV rows Polars scans for schema inference. |
DATASETPEEK_S3_DOWNLOAD_TIMEOUT_SECONDS |
30 |
Timeout for S3-compatible object downloads. |
Run the test suite:
make testRun tests plus bytecode compilation checks:
make checkapp/
main.py Robyn app setup and route registration
routes/ HTTP handlers
services/ File loading, profiling, heuristics, and view-model assembly
templates/ Jinja templates
static/ CSS and browser assets
app/img/ Logo and icon source assets
tests/ Route, service, delimiter, and storage tests
docs/PRD.md Product requirements
docs/branch-protection-plan.md
Terraform-managed GitHub branch protection policy
infra/github/ Terraform for GitHub repository settings
agents.md Repository guidance for coding agents
main.py Thin root launcher
Makefile Common developer commands
make install # sync dependencies
make run # start the Robyn app
make test # run pytest
make check # run tests and py_compile
make clean # remove pytest and Python cache filesGitHub branch protection is managed with Terraform in infra/github.
The current solo-maintainer policy requires PRs and the test CI check for master, without requiring a second approving reviewer.
See docs/branch-protection-plan.md for the policy and apply workflow.
render.yaml configures DatasetPeek as a single Render web service on Render's free plan.
- Health check:
/health - Build command:
pip install uv && uv sync --locked - Start command:
uv run python main.py - Python version:
3.12.10
Operational assumptions for this deployment:
- Uploads are processed in-request; preview sample sets are embedded in the rendered response.
- Restart, redeploy, crash, or free-tier spin-down clears any in-flight request state.
- The service should stay at a single instance unless upload state is moved out of memory.
- Render free web services spin down after 15 minutes of inactivity, so the first request after idle can take about a minute to recover.
- Free web services do not support persistent disks or scaling beyond a single instance.
- Keep uploads modest in size. The app warns above 50 MB and rejects uploads above 100 MB.
- Configure S3-compatible credentials in Render environment variables; Render does not read your local
.envfile.