ToolCabana

Dokument zu strukturiertem JSON

Exportiere PDF-Textblöcke mit Seitenzahlen und Begrenzungsrahmen.

Favorites are saved in this browser. Find them in My favorites.

Dokument zu strukturiertem JSON

Your input stays in this browser unless stated otherwise
Clear your input, files and result, and restore the default settings.

Drop your files here

or choose files from your device

So verwendest du Dokument zu strukturiertem JSON

Document to structured JSON exports the text of a PDF as JSON, listing every page with its width and height and each text block with its text, x and y position, width and height. Blocks are sorted top to bottom and left to right. You can process up to 100 pages, and the file downloads as document.json.

  1. Wähle eine Quelldatei aus. Wenn das Tool mehrere Dateien akzeptiert, ordne sie in der gewünschten Reihenfolge an.
  2. Lege die maximale Seitenzahl fest.
  3. Führe „Dokument zu strukturiertem JSON“ aus, prüfe die Ausgabe und verwende dann die verfügbaren Bedienelemente zum Kopieren oder Herunterladen.

Was dieses Tool unterstützt

Extrahiert die vorhandene PDF-Textebene in Koordinatenreihenfolge. Scans erfordern zuerst OCR. Spalten und komplexe Layouts können manuelle Korrekturen erfordern; JSON behält Seitenzahlen und Begrenzungsrahmen bei.

Maximale Seitenzahl
Range: 1 to 100

Grenzen und Verarbeitung

Bildeingaben: bis zu 32 Megapixel, Stapel mit bis zu 10 Bildern. Audio-/Video-Uploads: 50 MB; Transkription: 120 Sekunden. Bearbeitete Videos: 20 Sekunden, 4 fps und 640 Pixel breit. Übersetzung: Englisch/Spanisch, 30.000 Zeichen; Untertitel: 200 Cues. Neuronale Sprachausgabe: 1.000 englische Zeichen. Modell-Downloads und Speicheranforderungen variieren je nach Gerät.

Beispielquelle
A useful tool makes everyday work easier. Review the results before sharing.

Häufige Fragen

What does the JSON output of a PDF look like?

It is an object with a pages array. Each page has page, width, height and blocks, and each block has text, x, y, width and height, so you can locate every text run on the page.

What units are the coordinates in the PDF JSON?

Positions and sizes use PDF points at a scale of 1, matching the page width and height in the same output. The y value is measured from the top of the page, so larger values sit lower down.

Can I get JSON from a scanned PDF?

Only after OCR. The tool reads the existing text layer and reports that no text layer was found when pages are images. Make the PDF searchable with OCR first, then export it.

Wo wird meine Eingabe verarbeitet?

Verarbeitet die Quelle in deinem Browser. Kopieren und Herunterladen sind ausdrückliche Aktionen; die Quelle wird nicht im Kontoverlauf gespeichert.

Welche Eingabelimits gelten?

Bildeingaben: bis zu 32 Megapixel, Stapel mit bis zu 10 Bildern. Audio-/Video-Uploads: 50 MB; Transkription: 120 Sekunden. Bearbeitete Videos: 20 Sekunden, 4 fps und 640 Pixel breit. Übersetzung: Englisch/Spanisch, 30.000 Zeichen; Untertitel: 200 Cues. Neuronale Sprachausgabe: 1.000 englische Zeichen. Modell-Downloads und Speicheranforderungen variieren je nach Gerät.