Paper2Audio text to speech (56 hrs/week free)

Paper2Audio text to speech (56 hrs/week free)

https://preview.redd.it/gzcuus9na6kh1.png?width=1500&format=png&auto=webp&s=3c6f81ea681800e835d443123155c1d180368eff

I’m Joe, the founder of Paper2Audio, a free text to speech reader app for listening to complex documents and books, with highly accurate narration and high-quality voices.  

What is Paper2Audio?
Most text to speech tools are not good at converting dense PDFs, research papers, textbooks and webpages to audio. Paper2Audio is built to turn complex text into accurate audio.  We also handle less complicated text, like standard EPUBs or plain text.

Paper2Audio is available on Android, iOS and on our website. Your documents automatically sync across devices.  We have a generous free plan for personal use (56 hours of audio generation per week, no credit card required for free plan), as well as a paid Plus subscription with higher audio and file/size limits ($20/month or $192 annually).  We also offer Enterprise plan options for teams.  

Promo offer
Our Plus plan is currently on sale until August 20 for $10/month for your first 3 months or $134.40 for your first year.  Subscribe at the sale rate from the pricing page of our website, or upgrade from the settings page of your account (discount applies automatically at checkout, no coupon code required).

File types: PDF, EPUB, DOCX, TEXT, webpages, plain text

Supported languages: Full support for English.  Beta support for Brazilian Portuguese, Chinese (Mandarin), French, German, Hindi, Italian, Marathi, Japanese, and Spanish.  

How is Paper2Audio different from other text to speech apps?

  1. Higher audio limits for our free plan (56 hours weekly audio generation) with high quality voices.
  2. Hyper-focus on accuracy:  Paper2Audio avoids reading things that usually make text to speech audio annoying, like repeated page numbers, headers, citations, footnotes, and unnecessary boilerplate. We clean up and normalize tricky text first, including math, Roman numerals, symbols, units, formulas, and other things that often sound wrong when read aloud by other text to speech services.
  3. Follow along with Reader View, our optimized version of the audio transcript: We reformat PDFs and other documents to fit your screen while including rich content like images and document formatting.  Use it to follow along with the audio, or to more easily read documents that are normally poorly formatted for small phone screens (like 2 column PDFs, tables and figures, etc).
    • Visual elements are included: “Tables, figures, images, and math appear inline and can be opened in a zoomable “figure view” pop-up. 
    • Single column view: Documents with multiple columns are displayed in a single column to improve readability on smaller screens.
    • Rich text formatting: We preserve the original formatting of your documents, including math, headings, lists, subscripts, and other inline styling, so you can skim, navigate and understand the document more quickly. Citations and footnotes are also included so that you know when an author is making a reference, but are only read aloud when needed to keep sentences intact.
  4. Summarizes visual elements like tables, math, and code or reads it aloud:  When adding a document, you can choose how tables, math, and code are narrated. Summary" (default) gives a concise summary of the item, while "Read as is" reads the content verbatim.  Or, you can skip narration for these elements during playback entirely. 
  5. Files automatically download to your phone for offline listening. 
  6. Publicly share your documents (they are private by default), including embedding the audio directly onto a webpage with an iframe snippet you can paste into WordPress, Substack, or any site that supports HTML. 
  7. Free pre-generated audiobooks that don’t count toward your audio generation quota (see the posts on our blog, with more coming regularly)
  8. Plus plan additional features: Audio mp4 file export for Plus users to edit or listen outside the Paper2Audio app, or export a processed document’s transcript as a Markdown file for use with other tools (web now, apps coming soon)

  

Any feedback or questions?
If you try Paper2Audio, I’d love to hear what works well, what doesn’t, and what feature or improvement would make the biggest difference for you. We are also working on adding more voices, so please let me know what additional narrator options you’d like to hear.

reddit.com
u/goldenjm — 2 days ago
▲ 13 r/ProductivityApps+1 crossposts

Paper2Audio updates: Making complex documents actually listenable (Free and paid plan options, sale until 8/20)

I’m Joe, the founder of Paper2Audio, a free text to speech reader app for listening to complex documents and books, with highly accurate narration and high-quality voices.  Our free plan allows 56 hours of audio generation per week.  Our paid version, Paper2Audio Plus, is currently on sale through August 20 at $10/month for your first 3 months or $134.40 for your first year.

A: What problem does Paper2Audio solve and what’s new with Paper2Audio since our last post?
Most text to speech tools are not good at converting dense PDFs, research papers, textbooks and webpages to audio. Paper2Audio is built to turn complex text into accurate audio.  We also handle less complicated text, like standard EPUBs or plain text.

Supported languages: Full support for English.  Beta support for Brazilian Portuguese, Chinese (Mandarin), French, German, Hindi, Italian, Japanese, and Spanish.  

Since my last r/iosapps post, we’ve added or improved:

  • You can now publicly share your documents, including embedding the audio directly onto a webpage with an iframe snippet you can paste into WordPress, Substack, or any site that supports HTML. 
  • Narration improvements: subscripts and superscripts spoken more naturally, better pronunciation for abbreviations and Roman numerals, and more accurate header removal.
  • Faster processing and downloads
  • Bookmarks to save your position while listening
  • Background audio support so music can keep playing while you listen
  • More languages (added Chinese-Mandarin, German, Hindi and Japanese in beta)
  • Better page rotation detection for scanned PDFs
  • Export a processed document’s transcript as a Markdown file for use with other tools (web now, apps coming soon, Plus plan only)  
  • Better pronunciation and accents for our British English voices
  • Playback highlighting moves more smoothly from word to word
  • Free pre-generated audiobooks that don’t count toward your audio generation quota (see the posts on our blog, with more coming regularly)

B: Why is Paper2Audio better than the top alternatives?

  1. Higher audio limits for our free plan (56 hours weekly audio generation) with high quality voices.

  2. Hyper-focus on accuracy:  Paper2Audio avoids reading things that usually make text to speech audio annoying, like repeated page numbers, headers, citations, footnotes, and unnecessary boilerplate. We clean up and normalize tricky text first, including math, Roman numerals, symbols, units, formulas, and other things that often sound wrong when read aloud by other text to speech services.

  3. Summarizes visual elements like tables, math, and code or reads it aloud:  When adding a document, you can choose how tables, math, and code are narrated. Summary" (default) gives a concise summary of the item, while "Read as is" reads the content verbatim.  Or, you can skip narration for these elements during playback entirely. 

  4. Follow along with Reader View, our optimized version of the audio transcript: We reformat PDFs and other documents to fit your screen while including rich content like images and document formatting.  Use it to follow along with the audio, or to more easily read documents that are normally poorly formatted for small phone screens (like 2 column PDFs, tables and figures, etc).

    • Visual elements are included: “Tables, figures, images, and math appear inline and can be opened in a zoomable “figure view” pop-up. 
    • Single column view: Documents with multiple columns are displayed in a single column to improve readability on smaller screens.
    • Rich text formatting: We preserve the original formatting of your documents, including math, headings, lists, subscripts, and other inline styling, so you can skim, navigate and understand the document more quickly. Citations and footnotes are also included so that you know when an author is making a reference, but are only read aloud when needed to keep sentences intact.
  5. Multiple playback modes for your document:  Choose to listen to your document in full, or to have us generate a long or short summary instead.  We recently improved summary length, structure, and scaling for longer documents.

 

C: Cost
Paper2Audio is available on iOS, Android and on our website. We have a generous free plan for personal use (56 hours of audio generation per week), as well as a paid Plus subscription with higher audio and file/size limits ($20/month or $192 annually).  We also offer Enterprise plan options for teams.  

Our Plus plan is currently on sale until August 20 for $10/month for your first 3 months or $134.40 for your first year.

Any feedback or questions?
If you try Paper2Audio, I’d love to hear what works well, what doesn’t, and what feature or improvement would make the biggest difference for you. We are also working on adding more narrators, so please let me know what additional voice types you’d like to hear.

u/goldenjm — 2 days ago

Rotated PDFs before OCR: Splitting rotation detection from correction

https://preview.redd.it/o3875bx3pp7h1.png?width=1268&format=png&auto=webp&s=50fec64a6e9db7eef65aebc1d0b970bb705b7ea6

I’m Joe, founder of Paper2Audio, a text-to-speech app that converts PDFs, articles, ebooks, and other documents into audio. We recently worked on a scanned-document preprocessing problem involving rotated PDF pages before OCR, and I thought the solution and tradeoffs might be relevant to others building document-vision pipelines.  You can read the full writeup here.

We found that about 5% of PDFs submitted to Paper2Audio are scanned documents, and about 5% of those files have incorrectly rotated pages. That created a real production problem for us: if a rotated page reaches OCR, the system can extract bad text, mess up reading order, or create incorrect bounding boxes, and those errors then flow directly into the generated audio and downstream document processing. So we needed a way to detect and correct rotated scanned pages before the main OCR/extraction step, without slowing down every document that users upload. 

We ended up splitting the fix into two stages:

1. Rotation check before OCR

We already rasterize a few pages early in our processing to detect the document’s primary language, so we reuse those images and send up to five sampled pages to a small vision model with a structured prompt: are any pages rotated, and if so by roughly 90, 180, or 270 degrees?

The rotation check (the “gate”) does not need to correct the document. It only needs to decide whether we should send the PDF to a slower correction path. That matters because most scans are already upright, so full correction on every file would waste latency and increase processing costs.

2. Page-level correction when flagged

If the gate flags the document as being rotated, a separate correction service processes the PDF page by page. For each page, it:

  • Turn the page into an image: Rasterize it with PyMuPDF at 2x zoom to increase visual detail.
  • Focus on the parts most likely to contain text: Instead of running OCR on the entire page in every possible orientation, we split the page into tiles and look for the densest ones. Text-heavy areas usually have lots of edges, so we use Canny edge density to find patches that are likely to contain useful text.
  • Reduce the number of rotations to test: A quick projection-profile check tells us whether the text lines appear mostly horizontal or vertical. That usually narrows the page down to either a “0 or 180 degrees” case or a “90 or 270 degrees” case, so we only need to test two orientations instead of four.
  • Use OCR to choose the correct orientation: For each candidate orientation, we sharpen the patch, run EasyOCR, and score the result based on OCR confidence. Text produces higher confidence scores when it is upright and lower scores when it is sideways or upside down, so the highest-scoring orientation is usually the right one.
  • Correct the PDF: Write the winning rotation into the PDF with set_rotation if it is non-zero. If a page fails to process, we leave it unchanged rather than guessing and potentially making the document worse.

We use EasyOCR for page rotation correction because its confidence scores are a useful signal for which orientation makes text most legible.

The final important design was what to do with uncertain rotations.  If a page cannot be corrected confidently, we leave it unchanged. A missed correction is usually less damaging than rotating a good page into the wrong orientation.

For our use case (small number of scanned documents), the lightweight routing gate makes page rotation detection and correction more practical. The gate and the correction system do not need to solve the same problem. The gate just needs to separate “probably safe to continue processing” from “worth spending more compute to fix rotation,” with false negatives treated as much more costly than false positives.  

I’d be interested to hear what other OCR/document-processing tools people have found useful for this kind of problem, especially for orientation detection, layout-aware preprocessing, or confidence scoring before the main extraction step. Are there better models or best practices for this task?

reddit.com
u/goldenjm — 2 months ago
▲ 5 r/ebooks

Free public domain audiobooks for your summer reading list

https://preview.redd.it/fw0j0w8xmj7h1.png?width=1536&format=png&auto=webp&s=5923f67f32e17fb66b7dc81af557664aaaac5bc7

Hi there, this is Joe, the founder of Paper2Audio, a text-to-speech app that converts PDFs, articles, ebooks, and other documents into audio (free to generate up to 56 hours of audio each week). 

Summer reading is more fun when you don’t have to sit still to do it, so we’re giving away audio access to six classic public domain books via Paper2Audio: three adventure picks and three romance picks! The collection includes shipwrecks, treasure hunts, wilderness survival, time-bending satire, sharp social comedy, slow-burn longing, and big emotional payoff. Listen while driving on a road trip, lying on the beach, weeding the garden, or just trying to make the most of a sunny afternoon.

  • A Connecticut Yankee in King Arthur’s Court by Mark Twain
  • Treasure Island by Robert Louis Stevenson
  • The Call of the Wild by Jack London
  • Pride and Prejudice by Jane Austen
  • The Blue Castle by L.M. Montgomery
  • Jane Eyre by Charlotte Brontë

Links to the audio files are available here.  These free titles are pre-converted to audio and won’t count toward your weekly audio generation limit. Just click any of the document links, then click “Save for later” at the top of the document to add it to your Paper2Audio listening queue.  Files are available for direct download for listening outside of the Paper2Audio app for Plus subscribers.  

Any thematic requests for future monthly free reads?

reddit.com
u/goldenjm — 2 months ago

We fixed a problem affecting ~0.25% of uploads. Here’s why it was still worth building.

https://preview.redd.it/9dvr29o6cj7h1.png?width=1268&format=png&auto=webp&s=90b7470c14250948111b55360bd81f8f87539b2d

I’m Joe, founder of Paper2Audio, a text-to-speech app that turns PDFs, articles, ebooks, and other documents into audio.  

We recently finished a project that sounds niche at first: fixing rotated scanned PDFs before running OCR and converting text to audio.  Around 5% of PDFs uploaded to Paper2Audio are scanned documents, and about 5% of those scans have pages rotated 90, 180, or 270 degrees. So we were dealing with roughly 0.25% of all PDF uploads.

At first glance, this is the kind of edge case we wouldn’t prioritize, but rotated scans are unusually painful for Paper2Audio because the failure turns into audio. If OCR reads a sideways page badly, it becomes garbled narration or hallucinated text that a user actually has to listen to. 

The strategic question was: how do we fix the edge case without slowing down the 99%+ of documents that do not need it?  To solve this problem, we split the solution into two steps: 

1. A quick “gate” check for all PDFs

We already rasterize a few pages early in the pipeline to detect the document’s primary language. We now reuse those images and send up to five sampled pages to a small OCR vision model (our “gate”) with a structured question: do any pages look rotated, and if so by roughly 90, 180, or 270 degrees?

That check does not need to fix the document. It only needs to decide whether the document is worth sending to a slower correction path.

2. A longer rotation correction path only when needed

If the gate flags the document as being rotated, we run page-level correction. The correction service rasterizes each page, looks for text-heavy regions, narrows the likely rotation angles, uses OCR to select the best page orientation, and writes the corrected rotation back into the PDF. If the system is uncertain, it leaves the page alone rather than guessing and potentially making things worse.

In this case, we decided that an issue affecting 0.25% of uploads is still worth fixing if:

  • the user experience is especially bad when it happens
  • the failure looks like your core product is broken
  • the problem happens before other important processing steps
  • you can build a routing layer so the fix only runs when needed

This also changed how we think about pipeline design. The expensive part does not always need to be optimized enough to run on every job. Sometimes the better strategy is to build a cheap “should we spend more compute?” decision in front of it.

We wrote up the more technical version here, including the OCR/image-processing details: https://www.paper2audio.com/posts/fixing-rotated-pdfs-tts

I’d love to hear how other founders think about this tradeoff: when do you decide an edge case is worth dedicated engineering time, even if it affects a very small percentage of users?

reddit.com
u/goldenjm — 2 months ago

Building text to speech tools: the hidden UX problem behind word highlighting

https://preview.redd.it/qchnme3zjj3h1.png?width=2816&format=png&auto=webp&s=a3112421e85eb4864cdbffe3a2e158ca05e66929

I’m Joe, the founder of www.Paper2Audio.com, a text to speech (TTS) tool that turns PDFs, research papers, ebooks, and web articles into audio.  We’re trying to share more of the behind-the-scenes product and engineering problems we run into while building Paper2Audio. 

We recently fixed a TTS problem that I thought might be interesting to other people building TTS, document parsing, and audio products.  The short version: word-level highlighting in an audio transcript gets complicated when the text shown to the user is not the same text read aloud by the TTS model.

Paper2Audio has “Reader View” where users can follow along with their document while listening.  We want Reader View to work as an enhanced audio transcript that preserves the original document as much as possible: equations should still look like equations, citations should still be visible but not read aloud, lists should keep their formatting, and the page should feel like a readable document rather than a raw TTS transcript.

In more complex documents like research papers or reports the displayed text might include math equations, HTML tags, Roman numerals, or other similar formatting. But the spoken text needs to be normalized first so it sounds right. For example, $x^2 + y^2 = r^2$ might be spoken as “x squared plus y squared equals r squared,” while the transcript UI still needs to highlight the math as it is being narrated.  

The mismatch between displayed and narrated words creates a timestamp problem for word highlighting. Our TTS model (Kokoro) gives us word-level timestamps for the spoken text, but our UI needs to highlight the formatted document text as it is being read aloud. A simple character-count mapping doesn’t work because the two strings can have different words, different punctuation, different lengths, and sometimes one visual token maps to many spoken words.

To solve this problem, we treat the spoken text and text displayed in Reader View as two separate versions of the same content, then apply the general alignment algorithm we developed between them.  After the TTS runs, we use matching words in both versions as anchors, then reconcile the mismatched regions between them.  Doing so allows us to display the original formatting in the audio transcript and make sure that the portion being read aloud is getting highlighted at the correct time in our Reader View.  You can read a longer version of this explanation in our blog post.  

The biggest lesson for us was that the maintainable solution was not adding more normalization rules.  We needed a new processing layer that explicitly connects the model-friendly representation to the user-friendly representation.  With this solution, when a citation is visible but skipped in speech, it does not get its own timestamp. When “Part III” is spoken as “Part 3,” it still lines up.  

I’d love feedback from anyone building in adjacent areas. Does this “two representations plus an alignment layer” approach match how you’d think about the problem, or would you try to preserve the mapping earlier in the pipeline? What other tradeoffs have you run into?

reddit.com
u/goldenjm — 3 months ago

Better word highlighting for complex text to speech documents

https://preview.redd.it/9af3otchwd2h1.png?width=2816&format=png&auto=webp&s=c1c3df477e5301b6b51467fff94c90dc0dec19e2

I’m Joe, the founder of Paper2Audio, a text to speech service that turns PDFs, research papers, ebooks, and web articles into audio, with a focus on accuracy for complex documents.

I really appreciate all of the outstanding feedback I’ve received from many members of this community.  We’ve made a ton of improvements to Paper2Audio based on your feature requests and issue reports.

One of the more interesting product/engineering problems we solved recently is how to handle word-level highlighting when the text spoken by the text to speech model is not the same text shown in our audio transcript UI.

In more complex documents like research papers or reports the displayed text might include math equations, HTML tags, Roman numerals, or other similar formatting. But the spoken text needs to be normalized first so it sounds right. For example, $x^2 + y^2 = r^2$ might be spoken as “x squared plus y squared equals r squared,” while the transcript UI still needs to highlight the math as it is being narrated.  We want users to still see the original rich document formatting, not the word-by-word audio transcript.

The mismatch between displayed and read aloud words creates a timestamp problem for word highlighting. Our TTS model (Kokoro) gives us word-level timestamps for the spoken text, but our UI needs to highlight the formatted document text as it is being read aloud. A simple character-count mapping doesn’t work because the two strings can have different words, different punctuation, different lengths, and sometimes one visual token maps to many spoken words.

To solve this problem, we treat the spoken text and text displayed in the audio transcript (our “Reader View”) as two separate versions of the same content, then apply the general alignment algorithm we developed between them.  After the TTS runs, we use matching words in both versions as anchors, then reconcile the mismatched regions between them.  Doing so allows us to display the original formatting in the audio transcript and make sure that the portion being read aloud is getting highlighted at the correct time in our Reader View.  Check out our blog post if you want more details.

With this solution, when a citation is visible but skipped in speech, it does not get its own timestamp. When “Part III” is spoken as “Part 3,” it still lines up.

Please let us know if you have any questions or feedback about this post.  We’re thinking about writing additional technical blog posts about different challenges we’ve encountered building Paper2Audio, so please feel free to request topics.

reddit.com
u/goldenjm — 3 months ago

Making text to speech word highlighting work for complex documents

https://preview.redd.it/a5dadaznr52h1.png?width=2816&format=png&auto=webp&s=6fa6ca14c57b1aba9b533603141bab3457a422a1

I’m Joe, the founder of Paper2Audio, a text to speech service that turns PDFs, research papers, ebooks, and web articles into audio, with a focus on accuracy for complex documents.

We’ve recently come up with a solution to a text to speech processing challenge: how to combine accurate text to speech pronunciation with a rich transcript view that maintains the formatting details of the original document, and keeps word-level highlighting accurate when the text shown to the user is not the same text spoken by the TTS model.

For example, in more complex documents like research papers or reports the displayed text might include math equations, HTML tags, markdown, Roman numerals, or other similar formatting. But the spoken text needs to be normalized first so it sounds right. For example, $x^2 + y^2 = r^2$ is read as  “x squared plus y squared equals r squared,” while the transcript highlights the math. 

We wrote up a blog post covering how we went about building a reconciliation algorithm that maps TTS word timestamps back onto the original formatted document.  Our solution is basically a translation layer after TTS. Our TTS model tells us when each word in the cleaned-up spoken text is said. We then line that back up with the richer document text users actually see. Instead of writing separate rules for equations, citations, formatting, and punctuation, we look for matching words in both versions and use them to keep the two texts synced and then word-level highlighting in the audio transcript (our “Reader View”) works properly. 

We were able to improve both the reading and the listening experience without changing the underlying TTS model itself. The audio output stays the same, but the post-processing layer lets us preserve rich document rendering, better pronunciation, and accurate highlighting at the same time.  

As far as we can tell, other text to speech services haven’t figured out how to solve this problem.  I would love feedback from people who have worked on TTS highlighting.  Does this general reconciliation approach match how you’d solve it?  Do you think there are any failure modes we should watch for?

reddit.com
u/goldenjm — 3 months ago
▲ 14 r/iosapps

Paper2Audio text to speech, now for reading documents too (Free and paid plan options)

https://preview.redd.it/nqnzyvvpuzzg1.jpg?width=5106&format=pjpg&auto=webp&s=9f0377c5f47e0956882438be922503ccc221bd87

I’m Joe, the founder of Paper2Audio, a free text to speech reader designed to help you get through long documents and books more efficiently, with high-accuracy narration for complicated material and high-quality voices.  Our free plan allows 56 hours of audio generation per week.  

A: What problem does Paper2Audio solve?
Most text to speech tools can handle simple documents, but not messy PDFs, research papers, reports, and books. Paper2Audio is built to turn complex documents into accurate audio. 

I posted in r/iosapps previously, with my most recent post here (4 months ago).  Since then, we’ve been working on a big improvement: making Paper2Audio better not just for listening to documents, but for reading along while you listen so it’s easier to stay focused, skim when needed, retain more from dense material, and get through your reading backlog more quickly. 

B: Why is Paper2Audio better than the top alternatives?

  1. Higher audio limits for our free plan (56 hours weekly audio generation) with high quality voices.
  2. Summarizes figures, tables, math, and even code into plain English so you’re not stuck with symbol-by-symbol or line-by-line narration.
  3. Hyper-focus on accuracy:  Paper2Audio avoids reading things that usually make text to speech audio annoying, like repeated page numbers, headers, footers, citations, footnotes, and unnecessary boilerplate. We clean up and normalize tricky text first, including math, code, abbreviations, Roman numerals, symbols, units, formulas, and other things that often sound wrong when read aloud by other text to speech services.
  4. Use Paper2Audio to read and follow along, not just listen:  With Reader View, our new method of reformatting PDFs and other documents to fit your screen while including rich content like images and document formatting:   
    • Visual elements are included in the transcript: If your document includes visual elements like tables, figures, images, or math, you can now see them directly in the audio transcript which makes it easier to follow along without losing your place or switching views. 
    • "Figure view" for visual elements: Click on any visual element to bring up the figure view pop up, then zoom and pan around the image for a more detailed view.
    • Single column view: Documents with multiple columns are displayed in a single column to improve readability on smaller screens.
    • Rich text formatting: We now preserve the original formatting of your documents, including math, headings, lists, subscripts, and other inline styling, so you can skim, navigate and understand the document more quickly. Citations are also included so that you know when an author is making a reference, but citation text is only read aloud when needed to keep sentences intact.

C: Cost
Paper2Audio is available on iOS, Android and on our website. We have a generous free plan for personal use (56 hours of audio generation per week), as well as a paid Plus subscription with higher audio and file/size limits ($20/month) for business users.

Any feedback or questions?
I’d love feedback on the new Reader View, the Paper2Audio listening experience, and the overall workflow.  What would make Paper2Audio your go-to tool for listening and reading your documents?

reddit.com
u/goldenjm — 3 months ago

https://preview.redd.it/vxjpe8qp1szg1.png?width=831&format=png&auto=webp&s=af7b5fb6968f69de56dd5eb9a11260dcc97d1224

I’m Joe, the founder of Paper2Audio, a free text to speech reader designed to help you get through long documents and books more efficiently, with high-accuracy narration for complicated material and high-quality voices.  Our free plan allows 56 hours of audio generation per week.  

I’ve posted previously in r/productivityapps (most recently 4 months ago here).  Since then, we’ve been working on a big improvement: making Paper2Audio better not just for listening to documents, but for reading along while you listen so it’s easier to stay focused, skim when needed, retain more from dense material, and get through your reading backlog more quickly.   

We are excited to announce Reader View, our new method of reformatting PDFs and other documents to fit your screen while including rich content like images and document formatting.  Reader View is the default when you open a document.

  • Visual elements are included in the transcript: If your document includes visual elements like tables, figures, images, or math, you can now see them directly in the audio transcript which makes it easier to follow along without losing your place or switching views. Any summary text for a visual element will only appear when it is being read to give you a more streamlined reading view.
  • "Figure view" for visual elements: Click on any visual element to bring up the figure view pop up, then zoom and pan around the image for a more detailed view.
  • Single column view: Documents with multiple columns are displayed in a single column to improve readability on smaller screens.
  • Rich text formatting: We now preserve the original formatting of your documents, including math, headings, lists, subscripts, and other inline styling, so you can skim, navigate and understand the document more quickly. Citations are also included so that you know when an author is making a reference, but citation text is only read aloud when needed to keep sentences intact.

What else has improved since our last post?

  • Better image and figure labeling and summaries: Significantly improved the detection and extraction of tables, figures and images, reduced repetitive summaries for the same figure or table, and improved summary accuracy in cases where an item has multiple sub-objects.  This makes technical documents faster to understand because the summaries are less repetitive and more useful.
  • Improved math detection and summaries: Better detection accuracy of math elements within documents so that non-math items are not incorrectly summarized.
  • Overhaul to citation processing: Complete update to our citation system, including changes to better distinguish between citations within a sentence vs. not for smoother listening without citations unnecessarily interrupting your focus.  In-sentence citations that would break a sentence are not removed.
  • More accurate Roman numeral processing: Major changes to overhaul detection and processing for Roman numerals so that they are pronounced correctly (there may still be issues with I and X since those characters often appear on their own in other non-Roman numeral contexts).
  • Faster overall document processing: About 10-20% faster processing for documents, (varies but complicated non-fiction PDFs will benefit most). 
  • Accessibility feature improvements for visually impaired users (better screen reader support and navigation)
  • Lots of bug fixes and minor UI improvements

Where do I get Paper2Audio and how much does it cost?
Paper2Audio is available on iOS, Android and on our website. We have a generous free plan for personal use (56 hours of audio generation per week), as well as a paid Plus subscription with higher audio and file/size limits ($20/month) for business users.

Any feedback or questions?
I’d love feedback on the new Reader View, the Paper2Audio listening experience, and the overall workflow. Is this something you’d use for PDFs, books, research papers, or work documents? What would you want improved?

reddit.com
u/goldenjm — 4 months ago