RAG Text Chunker
Split documents into token-sized overlapping chunks for embedding.
Source document
Drop a file here
Drop a .txt or .md file
Chunks
Estimating while the tokenizer loads
Largest chunk: 0 tokens.
- Paste a document to split it into chunks.
About the RAG Text Chunker
Prepare documents for retrieval-augmented generation: choose paragraph, sentence, Markdown-heading or fixed-window splitting, set chunk size and overlap in real tokens, preview every chunk with its token count and export ready-to-embed JSONL with source metadata.
Examples
512-token chunks
12 KB Markdown guideOutput
7 chunks · 3,180 tokens · $0.00006 to embedKeyboard shortcuts
- Copy the main outputCtrl / ⌘ + Shift + C
- Download the resultCtrl / ⌘ + S
- Share this toolCtrl / ⌘ + Shift + S
- Reset the inputsAlt + R
- Open the tool search paletteCtrl / ⌘ + K
Related tools
AI Dataset Deduplicator
AI & LLM
Remove exact and near-duplicate rows from fine-tuning datasets.
AI Dataset Statistics Analyzer
AI & LLM
Profile a JSONL training set: token stats, roles, duplicates and cost.
AI PII & Secret Redactor
AI & LLM
Strip emails, keys, tokens and personal data before prompting an LLM.
AI Token Counter
AI & LLM
Count exact LLM tokens for GPT, Claude and Gemini prompts.
AI Tool Schema Builder
AI & LLM
Build function-calling schemas for OpenAI, Claude, Gemini, MCP and the AI SDK.
AI Tool Schema Converter
AI & LLM
Convert function-calling schemas between OpenAI, Anthropic, Gemini and MCP.
Frequently asked questions
Version 1.0.0 · Updated 2026-08-14 · Runs entirely in your browser