news.volyx.in

Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

562 points by Jeannen · 1207 days ago · 281 comments on HN

Article summary

The author proposes creating an 'ai.txt' file to provide information about a website, such as its purpose and author, to help AI crawlers understand the site's content without having to infer it. This file could be useful if a website ends up in a training dataset. The author suggests that this could be helpful for AI website crawlers. The idea is to provide a concise description of the page content in a way that is easily accessible to AI models.

Main themes

  • AI training data
  • copyright and legislation
  • website metadata
  • robots.txt and ai.txt
  • AI model behavior
  • information security
  • metadata standards

What commenters say

  • The proposal for 'ai.txt' is unnecessary because existing standards like 'robots.txt' can be extended to serve the same purpose.
  • The 'ai.txt' file could give a false sense of security and is not a reliable way to control how AI models use website content.
  • The main purpose of 'ai.txt' is to provide helpful information to AI crawlers, not to set boundaries for them.
  • The idea of 'ai.txt' is flawed because it assumes that AI models will respect the file's contents, which may not be the case.
  • The use of 'ai.txt' could potentially become an adversarial vector for attempts to trick AI models into misclassifying information.
  • Existing metadata standards, such as HTML tags, already provide a way to include structured metadata in pages, making 'ai.txt' redundant.
  • The proposal for 'ai.txt' highlights the need for clearer guidelines and legislation around AI training data and copyright.
  • The 'ai.txt' file could be used to signal the intention of the copyright holder and provide a clear indication of what content is allowed to be used for AI training.