news.volyx.in

Anyone got a contact at OpenAI. They have a spider problem (mailman.nanog.org)

743 points by speckx · 856 days ago · 375 comments on HN

Article summary

The article appears to be about a website with a large number of subdomains, each with one page, designed to waste the time of web crawlers. The website's owner claims that OpenAI's crawler, GPTBot, is not respecting the website's robots.txt file and is crawling the site despite being disallowed. The owner is informing OpenAI of this issue, which may be impacting the performance of the website for legitimate users. The website's purpose is to test and waste the resources of crawlers.

Main themes

  • Web Crawling
  • Robots.txt
  • AI Training
  • Web Scraping
  • Crawler Evasion

What commenters say

  • OpenAI's GPTBot is not respecting the website's robots.txt file, which is a standard protocol for controlling crawler access.
  • The website's design is intended to waste the time and resources of web crawlers, and OpenAI's crawler is falling into this trap.
  • Some commenters believe that OpenAI should be aware of the website's robots.txt file and respect its directives, while others think that the company may be intentionally ignoring it.
  • The legality of web scraping is disputed, and some commenters argue that it is legal to scrape public websites, while others point out that terms of service agreements can prohibit scraping.
  • The cost of training AI models like GPT-3 and GPT-4 is significant, but the claim of 'millions of dollars per second' is likely an exaggeration.
  • The website's owner may be trying to help OpenAI improve its crawler by pointing out its flaws, rather than trying to harm the company.
  • Some commenters are skeptical of the website's true intentions and believe that it may be trying to lure OpenAI into a trap or create a denial-of-service attack.