ToolCabana

Documento a JSON estructurado

Exporta bloques de texto de PDF con números de página y recuadros.

Favorites are saved in this browser. Find them in My favorites.

Documento a JSON estructurado

Your input stays in this browser unless stated otherwise
Clear your input, files and result, and restore the default settings.

Drop your files here

or choose files from your device

Cómo usar Documento a JSON estructurado

Document to structured JSON exports the text of a PDF as JSON, listing every page with its width and height and each text block with its text, x and y position, width and height. Blocks are sorted top to bottom and left to right. You can process up to 100 pages, and the file downloads as document.json.

  1. Elige un archivo de origen. Si la herramienta acepta varios archivos, ordénalos como quieras.
  2. Configura las páginas máximas.
  3. Ejecuta la conversión de documento a JSON estructurado, revisa la salida y usa los controles disponibles para copiar o descargar.

Qué admite esta herramienta

Extrae la capa de texto existente del PDF en orden de coordenadas. Los escaneos requieren OCR primero. Las columnas y diseños complejos pueden necesitar correcciones manuales; JSON conserva números de página y recuadros.

Páginas máximas
Range: 1 to 100

Límites y procesamiento

Entradas de imagen: hasta 32 megapíxeles, lotes de hasta 10 imágenes. Subidas de audio/vídeo: 50 MB; transcripción: 120 segundos. Vídeos editados: 20 segundos, 4 fps y 640 píxeles de ancho. Traducción: inglés/español, 30,000 caracteres; subtítulos: 200 entradas. Voz neuronal: 1,000 caracteres en inglés. Las descargas de modelos y los requisitos de memoria varían según el dispositivo.

Fuente de ejemplo
A useful tool makes everyday work easier. Review the results before sharing.

Preguntas frecuentes

What does the JSON output of a PDF look like?

It is an object with a pages array. Each page has page, width, height and blocks, and each block has text, x, y, width and height, so you can locate every text run on the page.

What units are the coordinates in the PDF JSON?

Positions and sizes use PDF points at a scale of 1, matching the page width and height in the same output. The y value is measured from the top of the page, so larger values sit lower down.

Can I get JSON from a scanned PDF?

Only after OCR. The tool reads the existing text layer and reports that no text layer was found when pages are images. Make the PDF searchable with OCR first, then export it.

¿Dónde se procesa mi entrada?

Procesa la fuente en tu navegador. Copiar y descargar son acciones explícitas; la fuente no se guarda en el historial de la cuenta.

¿Cuáles son los límites de entrada?

Entradas de imagen: hasta 32 megapíxeles, lotes de hasta 10 imágenes. Subidas de audio/vídeo: 50 MB; transcripción: 120 segundos. Vídeos editados: 20 segundos, 4 fps y 640 píxeles de ancho. Traducción: inglés/español, 30,000 caracteres; subtítulos: 200 entradas. Voz neuronal: 1,000 caracteres en inglés. Las descargas de modelos y los requisitos de memoria varían según el dispositivo.