news.volyx.in

What's so hard about PDF text extraction? (filingdb.com)

733 points by maest · 2411 days ago · 342 comments on HN

Article summary

The article discusses the challenges of extracting text from PDFs, but its content is not available. Commenters share their experiences and opinions on the matter. Some suggest using OCR techniques, while others propose embedding data within the PDF itself. The discussion highlights the difficulties of working with PDFs and the need for better tools and methods for text extraction.

Main themes

  • PDF text extraction
  • OCR techniques
  • Embedded data
  • PDF forms
  • Automation
  • Data extraction tools
  • Error reduction

What commenters say

  • Using OCR techniques can be effective for extracting text from PDFs, but it may not always produce accurate results.
  • Embedding data within the PDF itself can make text extraction easier and more reliable.
  • Some commenters argue that breaking the browser's back navigation is not an effective way to impede automatic HTML text extraction.
  • Others propose using dedicated parsers or services to extract data from PDFs and fill out forms.
  • There is a need for better tools and methods for text extraction from PDFs, as current solutions can be time-consuming and prone to errors.
  • Some services, such as hellosign.com, can already solve the problem of turning a PDF into a web form and collecting signatures.
  • The use of PDF forms with built-in validation can simplify the process of extracting data and reduce errors.
  • Manual entry of data from PDFs can be tedious and error-prone, and automation can help improve efficiency and accuracy.