ToolCabana

Divisor de fragmentos de documento

Divide texto en fragmentos solapados de caracteres Unicode.

Favorites are saved in this browser. Find them in My favorites.

Divisor de fragmentos de documento

Your input stays in this browser unless stated otherwise
Clear your input, files and result, and restore the default settings.
Load sample text to explore what this tool can do. This replaces your current input.
0 charactersClear the source text.

Cómo usar Divisor de fragmentos de documento

Document chunk splitter cuts text into overlapping chunks of a fixed number of Unicode characters, for example to prepare content for retrieval or embedding. Chunk size ranges from 100 to 10,000 characters with an overlap smaller than the chunk size. The result is a JSON array listing each chunk's id, start and end positions and text.

  1. Pega la fuente en el editor de entrada o carga el ejemplo integrado.
  2. Configura los caracteres de fragmento y de solapamiento.
  3. Ejecuta el divisor de fragmentos de documento, revisa la salida y usa los controles disponibles para copiar o descargar.

Qué admite esta herramienta

Divide texto en fragmentos solapados de caracteres Unicode.

Caracteres de fragmento
Range: 100 to 10000
Caracteres de solapamiento
Range: 0 to 9999

Límites y procesamiento

Se aplican los límites de archivos y texto del navegador, junto con los límites mostrados en Ajustes. El procesamiento de imágenes está limitado a 32 megapíxeles. La calidad del OCR depende de la resolución y el diseño de la fuente.

Fuente de ejemplo
A useful tool makes everyday work easier. Review the results before sharing.

Preguntas frecuentes

What chunk size and overlap should I use for RAG?

The defaults are 1,200 characters with 150 characters of overlap. Size can be 100 to 10,000, and overlap can be 0 or more but must stay smaller than the chunk size.

Does the chunk splitter count tokens?

No. Chunks are measured in Unicode characters, counting each code point once, and the character count is not the same as a model token count.

Does the document chunk splitter break at sentences?

No. Chunks are cut at fixed character positions, with each new chunk starting the overlap distance before the previous one ended. Sentences and words can be split between chunks.

¿Dónde se procesa mi entrada?

Procesa la fuente en tu navegador. Copiar y descargar son acciones explícitas; la fuente no se guarda en el historial de la cuenta.

¿Cuáles son los límites de entrada?

Se aplican los límites de archivos y texto del navegador, junto con los límites mostrados en Ajustes. El procesamiento de imágenes está limitado a 32 megapíxeles. La calidad del OCR depende de la resolución y el diseño de la fuente.