Document chunk splitter
Split text into overlapping Unicode character chunks.
Document chunk splitter
วิธีใช้ Document chunk splitter
Document chunk splitter cuts text into overlapping chunks of a fixed number of Unicode characters, for example to prepare content for retrieval or embedding. Chunk size ranges from 100 to 10,000 characters with an overlap smaller than the chunk size. The result is a JSON array listing each chunk's id, start and end positions and text.
- Paste the source into the input editor, or load the built-in example.
- Set chunk characters, overlap characters.
- Run document chunk splitter, review the output, then use the available copy or download controls.
ความสามารถของเครื่องมือนี้
Split text into overlapping Unicode character chunks.
- Chunk characters
- Range: 100 to 10000
- Overlap characters
- Range: 0 to 9999
ข้อจำกัดและการประมวลผล
Browser file and text limits apply, along with the bounds shown in Settings. Image processing is capped at 32 megapixels. OCR quality depends on source resolution and layout.
ตัวอย่างข้อมูลนำเข้า
A useful tool makes everyday work easier. Review the results before sharing.
คำถามที่พบบ่อย
What chunk size and overlap should I use for RAG?
The defaults are 1,200 characters with 150 characters of overlap. Size can be 100 to 10,000, and overlap can be 0 or more but must stay smaller than the chunk size.
Does the chunk splitter count tokens?
No. Chunks are measured in Unicode characters, counting each code point once, and the character count is not the same as a model token count.
Does the document chunk splitter break at sentences?
No. Chunks are cut at fixed character positions, with each new chunk starting the overlap distance before the previous one ended. Sentences and words can be split between chunks.
Where is my input processed?
Processes source in your browser. Copy and download are explicit actions; source is not saved to account history.
What are the input limits?
Browser file and text limits apply, along with the bounds shown in Settings. Image processing is capped at 32 megapixels. OCR quality depends on source resolution and layout.