ToolCabana

Documento in JSON strutturato

Esporta i blocchi di testo PDF con numeri di pagina e riquadri di delimitazione.

Favorites are saved in this browser. Find them in My favorites.

Documento in JSON strutturato

Your input stays in this browser unless stated otherwise
Clear your input, files and result, and restore the default settings.

Drop your files here

or choose files from your device

Come usare Documento in JSON strutturato

Document to structured JSON exports the text of a PDF as JSON, listing every page with its width and height and each text block with its text, x and y position, width and height. Blocks are sorted top to bottom and left to right. You can process up to 100 pages, and the file downloads as document.json.

  1. Scegli un file di origine. Se lo strumento accetta più file, disponili nell'ordine desiderato.
  2. Imposta il numero massimo di pagine.
  3. Esegui Documento in JSON strutturato, esamina il risultato, poi usa i controlli disponibili per copiare o scaricare.

Cosa supporta questo strumento

Estrae il livello di testo PDF esistente in ordine di coordinate. Le scansioni richiedono prima l'OCR. Colonne e layout complessi possono richiedere correzioni manuali; JSON conserva numeri di pagina e riquadri di delimitazione.

Numero massimo di pagine
Range: 1 to 100

Limiti ed elaborazione

Input immagine: fino a 32 megapixel, batch fino a 10 immagini. Caricamenti audio/video: 50 MB; trascrizione: 120 secondi. Video modificati: 20 secondi, 4 fps e 640 pixel di larghezza. Traduzione: inglese/spagnolo, 30.000 caratteri; sottotitoli: 200 cue. Sintesi vocale neurale: 1.000 caratteri inglesi. Download dei modelli e requisiti di memoria variano in base al dispositivo.

Origine di esempio
A useful tool makes everyday work easier. Review the results before sharing.

Domande frequenti

What does the JSON output of a PDF look like?

It is an object with a pages array. Each page has page, width, height and blocks, and each block has text, x, y, width and height, so you can locate every text run on the page.

What units are the coordinates in the PDF JSON?

Positions and sizes use PDF points at a scale of 1, matching the page width and height in the same output. The y value is measured from the top of the page, so larger values sit lower down.

Can I get JSON from a scanned PDF?

Only after OCR. The tool reads the existing text layer and reports that no text layer was found when pages are images. Make the PDF searchable with OCR first, then export it.

Dove viene elaborato il mio input?

Elabora i dati di origine nel tuo browser. Copia e download sono azioni esplicite; l'origine non viene salvata nella cronologia dell'account.

Quali sono i limiti di input?

Input immagine: fino a 32 megapixel, batch fino a 10 immagini. Caricamenti audio/video: 50 MB; trascrizione: 120 secondi. Video modificati: 20 secondi, 4 fps e 640 pixel di larghezza. Traduzione: inglese/spagnolo, 30.000 caratteri; sottotitoli: 200 cue. Sintesi vocale neurale: 1.000 caratteri inglesi. Download dei modelli e requisiti di memoria variano in base al dispositivo.