news.volyx.in

TimeCapsuleLLM: LLM trained only on data from 1800-1875 (github.com)

737 points by admp · 191 days ago · 314 comments on HN

Article summary

A language model called TimeCapsuleLLM has been trained exclusively on data from 1800-1875 to reduce modern bias and emulate the voice, vocabulary, and worldview of that era. The model has been trained on a dataset of 90GB of text from London during that time period. The goal of the project is to create a language model that can reason exclusively using knowledge from that time period. The model has shown promising results, with outputs that reflect the linguistic style and historical context of the time period.

Main themes

  • Historical language models
  • Reducing modern bias
  • Emulating historical voice and vocabulary
  • Language and cognition
  • Artificial intelligence and machine learning

What commenters say

  • Training a language model on pre-1900 data and prompting it about modern concepts like quantum mechanics could be a interesting test of its capabilities.
  • Some commenters argue that language models are incapable of true thought or reasoning, and are limited to predicting tokens based on patterns in the data they were trained on.
  • Others propose that language may play a more central role in human cognition than previously thought, and that manipulating tokens of language could be a key aspect of human thought.
  • There is disagreement about whether language models can truly be said to be 'thinking' or 'reasoning', with some arguing that they are simply imitating human-like language patterns.
  • The relationship between language, cognition, and thought is not fully understood, and more research is needed to determine the capabilities and limitations of language models.
  • Some commenters argue that the fact that language models have not produced novel scientific concepts or theories is evidence that they are not truly capable of thought or reasoning.
  • Others suggest that the fact that humans are able to think and reason in ways that language models currently cannot may be due to the complexity of human cognition, rather than any inherent limitation of language models.