news.volyx.in

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks (z.ai)

484 points by CuriouslyC · 160 days ago · 520 comments on HN

Article summary

The article appears to discuss the release of GLM-5, a new model available on the chat endpoint of a company's website, with some users experiencing issues accessing it due to traffic and scaling problems. The model is not yet available for all users, with some plans not including GLM-5 quota. The company has announced that they will gradually expand the scope and enable more users to experience and use GLM-5. Details about the model's capabilities and pricing are not fully clear from the discussion.

Main themes

  • GLM-5 model release
  • Access and scaling issues
  • Pricing and plans
  • Model capabilities
  • Self-hosting and local inference
  • Hardware requirements

What commenters say

  • The new GLM-5 model is not yet widely available due to traffic and scaling issues, with some users experiencing errors when trying to access it.
  • Some users believe that self-hosting and local inference are important for owning one's intelligence and maintaining privacy, despite the high costs and hardware requirements.
  • Others argue that cloud inference is more practical and cost-effective, with some providers offering cheap APIs and flexible pricing plans.
  • There is disagreement about the feasibility of running large models like GLM-5 on consumer-grade hardware, with some arguing that Apple devices offer sufficient memory bandwidth and others claiming that they are inferior to Nvidia hardware for certain tasks.
  • The cost of running models locally can be high, with some users estimating that it would take years to break even on the cost of hardware, and others pointing out that electricity costs and other factors must be considered.
  • Some users are exploring alternative options, such as using open-source models and running them on headless Linux boxes or other custom hardware setups.