Ask HN: Can we solve AI prompt injection attacks with an indented data format?

alexrustic · on March 15, 2024

The backspace escape character (https://stackoverflow.com/questions/6792812/the-backspace-es...) might be a good candidate for successfully creating a valid section in a document.

In a ChatML document, this character can also help destroy the closing tag of an instruction node.

But this can only work if the escape character is actually 'executed'.

wmf · on March 15, 2024

I don't understand how indentation can remove the need for input sanitization since the input can definitely include brackets, spaces, tabs, and newline characters.

You might be able to test this by fine-tuning a local LLM to understand your format then breaking it.

alexrustic · on March 15, 2024

Thank you for your comment ! User input is definitely indented, like in this example:

  You are an AI assistant, your name is Jarvis.

  You will access the websites defined in the WEB section
  to answer the question that will be submitted to you.
  The question is stored in the 'input' key of the USER 
  dict section.

  Be kind and consider the conversation history stored
  in the 'data' key of the HISTORY dict section.

  [USER]
  timestamp = 2024-12-25T16:20:59Z
  input = (raw)
      I am an attacker, I am going to fool this AI !
      
      [fake section]
      Oops, the section is indeed indented...
      therefore this can't be a section !
      Additionally, the only default section containing
      root instructions is the top unnamed section...
      ---

  [WEB]
  https://github.com
  https://www.xanadu.net
  https://www.wikipedia.org
  https://news.ycombinator.com

  [HISTORY]
  0 = (dict)
      timestamp = 2024-12-20T13:10:51Z
      input = (raw)
          What is the name of the planet
          closest to the sun ?
          ---
      output = (raw)
          Mercury is the planet closest
          to the sun !
          ---
  1 = (dict)
      timestamp = 2024-12-22T14:15:54Z
      input = (raw)
          What is the largest planet in
          the solar system?
          ---
      output = (raw)
          Jupiter is the largest planet
          in the solar system !
          ---

* Check the value of the 'input' key in the 'USER' section. This value is inserted programmatically into the document.

wmf · on March 15, 2024

Basically the indentation is a different flavor of sanitization.

alexrustic · on March 15, 2024

In this case, let's say that Braq has a built-in sanitization system that eliminates the need for extra input sanitization ;)