news.volyx.in

A Facebook crawler was making 7M requests per day to my stupid website (coding.napolux.com)

1074 points by napolux · 2307 days ago · 407 comments on HN

Article summary

A website received 7 million requests per day from a Facebook crawler, causing concern for the website owner. The owner, a staff engineer with 20 years of experience, shared their story and provided IP addresses as evidence. The requests were likely made to gather information for link previews when users share URLs on Facebook. The website owner's experience sparked a discussion about web crawlers and HTTP protocols.

Main themes

  • Facebook crawler
  • HTTP protocols
  • web scraping
  • link previewing
  • idempotent requests
  • safe requests

What commenters say

  • Some commenters questioned whether the requests actually came from Facebook or if they were spoofed by a malicious actor.
  • Others argued that it's possible for IP addresses to be faked, but establishing a TCP connection makes it harder to fake the source IP.
  • Commenters discussed the importance of following HTTP protocol standards, particularly regarding idempotent and safe requests.
  • There was disagreement over whether Facebook's crawler could be misbehaving, with some arguing that it's unlikely and others saying it's possible.
  • Some argued that using GET requests for actions that modify system state is bad practice and can cause problems.
  • Others pointed out that Facebook's crawler is likely used to gather metadata for link previews, and that this is a common practice among search engines and social media platforms.
  • A few commenters expressed frustration with the practice of link previewing and the noise it can generate in conversations.
  • Some discussed the distinction between idempotent and safe requests, and how they are important for preventing unintended actions.