Every page is read
Not just the text, but the layout: where the headline is, which caption belongs to which photo, where a new article starts and where it continues on the next page.
You supply the PDF of an edition. The tool reads every page, recognises headlines, standfirsts, body text, captions, boxes and photos, and turns them into separate articles. Ready to publish, archive, or hand over to your own system.
Most publishers have already laid out their edition perfectly in a PDF, and then have to build it a second time by hand for online. Retyping, copying and pasting, pulling out photos, hunting down captions. Every single issue.
A PDF does not know what a headline is and what body text is. It only knows where each letter sits. That is why plain text extraction gives you a soup that nobody can use.
Not just the text, but the layout: where the headline is, which caption belongs to which photo, where a new article starts and where it continues on the next page.
Headline, kicker, section label, standfirst, body text, crosshead, pull quote, caption, byline, table, advertisement. So your system knows what to do with each piece of text.
Images are extracted from the PDF as they are, with the caption attached. Not screenshots of pages.
Every publication has its own layout. After the first issue the tool records the rules for your title and applies them to everything after.
Publish the edition straight away as a digital issue where readers page through and open articles in readable form, on phone and desktop.
Export every edition as a ZIP with the articles in HTML and the data in JSON, photos included. Feed it into your own CMS.
Publishers of trade journals, association magazines, consumer magazines and local newspapers who want their print edition online as well, without turning it into an editorial chore.
Also suited to publishing groups with several titles: each publication keeps its own rules while the workflow stays the same.