news.volyx.in

Bing ChatGPT image jailbreak (twitter.com)

464 points by tomduncalf · 1058 days ago · 236 comments on HN

Article summary

Denis Shiryaev claims to have successfully used Bing ChatGPT to read a captcha by using prompt-visual engineering. This was done by manipulating the prompt to get the AI to quote the captcha. The method was later patched by Bing. The discussion around this topic explores the limitations and potential of AI models like ChatGPT.

Main themes

  • LLM limitations
  • AI security
  • Social engineering
  • Guardrails and control
  • Human-AI parallels
  • Prompt injection vulnerabilities

What commenters say

  • Some argue that if LLMs were truly intelligent, they could be stopped from doing something simply by being told not to, without needing complex guardrails.
  • Others counter that humans can also be vulnerable to social engineering, and that this does not necessarily prove or disprove the intelligence of LLMs.
  • The ability to bypass LLM guardrails through creative prompts is seen as a limitation of current AI models.
  • Implementing effective guardrails for LLMs is a challenging task, and some argue that current methods are not sufficient.
  • The concept of jailbreaking LLMs highlights the need for more sophisticated strategies to control their behavior.
  • Some commenters draw parallels between the vulnerabilities of LLMs and those of humans, suggesting that both can be manipulated through careful wording or social engineering.
  • The effectiveness of prompt-based guardrails for LLMs is disputed, with some arguing that they can be bypassed too easily.
  • The development of more robust and secure LLMs will require addressing the challenges of prompt injection and social engineering.