AI Dataset Statistics Analyzer
Profile a JSONL training set: token stats, roles, duplicates and cost.
Dataset (JSONL)
Drop a file here
or choose a file — it never leaves your device
Report
Token length distribution
- 0–20
- 2–40
- 4–60
- 6–81
- 8–100
- 10–121
- 12–140
- 14–161
- 16–180
- 18–200
Role distribution
- user2
- assistant2
- system1
- prompt1
- completion1
About the AI Dataset Statistics Analyzer
Analyse a fine-tuning or evaluation dataset in one pass: record count, total and average tokens, min, max and p50/p90/p99 percentiles, token length histogram, role distribution, empty and duplicate rows, records that exceed the model context and the input cost of a single pass.
Examples
Chat dataset
3 JSONL recordsOutput
avg 24 tokens, p90 38, 0 duplicatesKeyboard shortcuts
- Copy the main outputCtrl / ⌘ + Shift + C
- Download the resultCtrl / ⌘ + S
- Share this toolCtrl / ⌘ + Shift + S
- Reset the inputsAlt + R
- Open the tool search paletteCtrl / ⌘ + K
Related tools
AI Dataset Deduplicator
AI & LLM
Remove exact and near-duplicate rows from fine-tuning datasets.
Chat Dataset Converter
AI & LLM
Convert between OpenAI chat, ShareGPT, Alpaca and prompt/completion.
AI PII & Secret Redactor
AI & LLM
Strip emails, keys, tokens and personal data before prompting an LLM.
AI Token Counter
AI & LLM
Count exact LLM tokens for GPT, Claude and Gemini prompts.
AI Tool Schema Builder
AI & LLM
Build function-calling schemas for OpenAI, Claude, Gemini, MCP and the AI SDK.
AI Tool Schema Converter
AI & LLM
Convert function-calling schemas between OpenAI, Anthropic, Gemini and MCP.
Frequently asked questions
Version 1.0.0 · Updated 2026-08-15 · Runs entirely in your browser