news.volyx.in

Perplexity AI is lying about their user agent (rknight.me)

626 points by cdme · 789 days ago · 533 comments on HN

Article summary

The author discovered that Perplexity AI is ignoring robots.txt and using a generic user agent to scrape content from their website, despite claiming to respect robots.txt. The author tested this by blocking PerplexityBot in their robots.txt and server configuration, but Perplexity AI was still able to access and summarize their content. The author found that Perplexity AI is using a headless browser to scrape content, which does not send the correct user agent string. This allows Perplexity AI to bypass blocking measures.

Main themes

  • AI content scraping
  • robots.txt ignoring
  • user agent spoofing
  • web scraping ethics
  • content licensing
  • AI company practices

What commenters say

  • Perplexity AI's actions are a clear violation of web scraping ethics and content licensing agreements.
  • The company's use of a generic user agent to scrape content is a deliberate attempt to bypass blocking measures.
  • Some argue that AI companies should be allowed to scrape content for the purpose of training their models, as long as they follow certain guidelines.
  • Others believe that website owners have the right to control how their content is used and distributed, and that AI companies should respect robots.txt and other blocking measures.
  • There is a distinction between crawling for indexing purposes and browsing on behalf of a user, and AI companies should use different user agents for these different activities.
  • The use of headless browsers to scrape content raises concerns about the ability of website owners to control access to their content.
  • Some commenters argue that Perplexity AI's actions are not unique and that many AI companies engage in similar practices.
  • The issue of content scraping and licensing is complex and requires a nuanced approach that balances the needs of website owners with the needs of AI companies.