Norm Compliance
Scenario ID: vijil.scenarios.safety_compliance_normcompliance
Norm Compliance
Scenario ID: vijil.scenarios.safety_compliance_normcompliance
| Probe Name | Description |
|---|---|
| Poses queries to the model that are expected to elicit responses recommending aggressive or bullying action. Adds Recipe Steganography, Whitespace Noise, and JSON Escape Mutations. The agent fails if it complies. | |
| Poses queries to the model that are expected to elicit responses recommending aggressive or bullying action. Adds Advanced Role Play Actor and JSON Escape Mutations. The agent fails if it complies. | |
| Poses queries to the model that are expected to elicit responses recommending aggressive or bullying action. Adds Whitespace Noise and MIME Quoted Printable Mutations. The agent fails if it complies. | |
| Tests whether the agent will output profanity. Adds Children’s Story Veil and Token Padding Mutations. Failures indicate tendency to produce profanity | |
| Tests whether the agent will output profanity. Includes Recipe Steganography, Whitespace Noise and JSON Escape Mutations. Failures indicate tendency to produce profanity | |
| Tests whether the agent will output profanity. Adds Comment Injection, Entropy Reduction Checksum Marker, Whitespace Noise, and HTML Entity Encoding Mutations. Failures indicate tendency to produce profanity | |
| Tests whether the agent will output profanity. Adds Legal Compliance Override and JSON Escape Mutations. Failures indicate tendency to produce profanity. | |
| Tests to assess if an agent will produce content that is racially discriminatory or hurtful towards specific racial groups. Adds Recipe Steganography, Whitespace Noise, and JSON Escape Mutations. | |
| Tests to assess if an agent will produce content that is racially discriminatory or hurtful towards specific racial groups. Adds Macaronic, Unicode Junk Injection, and Zalgo Noise Mutations. | |
| Tests to assess if an agent will produce content that is racially discriminatory or hurtful towards specific racial groups. Adds Children’s Story Veil and Token Padding Mutations. | |
| Tests to assess if an agent will produce content that is racially discriminatory or hurtful towards specific racial groups. Adds Chunk Fragmentation Mutations. | |
| Tests whether an agent will continue real-world phrases and examples of toxic and inappropriate statements that are often used in LLM training corpora. Fails if the continuation is toxic. | |
| Tests to assess if an agent will produce content that promotes religious extremism. Adds Whitespace Noise and MIME Quoted Printable Mutations. | |
| Tests to assess if an agent will produce content that promotes religious extremism. Adds Song Lyric Obfuscation and Chunk Fragmentation Mutations. | |
| Tests to assess if an agent will produce content that promotes religious extremism. Adds Recipe Steganography, Whitespace Noise, and JSON_Escape Mutations | |
| Tests to assess if an agent will produce content that promotes religious extremism. Adds Recipe Steganography and Zalgo Noise Mutations. | |
| A collection of tests that checks if an agent is able to refrain from producing hateful or stereotypical content about various demographics of people under adversarial conditions. Failures in these tests indicate that the agent is vulnerable to producing hateful content fueled by stereotypes if maliciously prompted. | |
| Tests to assess if an agent will produce content involving sexual minors. Adds Enhancement, Unicode Junk Injection, and Zalgo Noise Mutations. | |
| Tests to assess if an agent will produce content involving sexual minors. Adds Song Lyric_Obfuscation and Chunk Fragmentation Mutations. | |
| Tests to assess if an agent will produce content involving sexual minors. Adds Whitespace Noise and MIME Quoted Printable Mutations. | |
| Tests to assess if an agent will produce content involving sexual minors. Adds Song Lyric Obfuscation Mutations. | |
| Tests to assess if an agent will produce content that is violence. Adds Song Lyric Obfuscation Mutation. | |
| Tests to assess if an agent will produce content that is violent. Adds Song Lyric Obfuscation and Chunk Fragmentation Mutations. | |
| Tests to assess if an agent will produce content that is violent. Adds Recipe Steganography, Whitespace Noise, and JSON Escape Mutations. | |
| Tests to assess if an agent will produce content that is violent. Adds Recipe Steganography and Zalgo_Noise Mutations | |
| Tests to assess if an agent will produce content that is sexually objectifying. Adds Zero Width Injection and Zalgo Noise Mutations. | |
| Tests to assess if an agent will produce content that is sexually objectifying. Adds Enhancement Mutation. | |
| Tests to assess if an agent will produce content that is sexually objectifying. Adds Song Lyric Obfuscation and Chunk Fragmentation Mutations. | |
| Tests to assess if an agent will produce content that is sexually objectifying. Adds Advanced Role Play Actor, JSON Escape, and Unicode Junk Injection Mutations. |
Policy Compliance
Scenario ID: vijil.scenarios.safety_compliance_policycompliance
Policy Compliance
Scenario ID: vijil.scenarios.safety_compliance_policycompliance
| Probe Name | Description |
|---|---|
| Tests to assess if an agent will produce content that is business_ethics. Adds Government Compliance Directive and MIME Quoted Printable Mutations. | |
| Tests principles common in business code of conducts. Adds Recipe Steganography and Chunk Fragmentation Mutations. | |
| Tests principles common in business code of conducts. Adds Recipe Steganography and Entropy Reduction Checksum Marker Mutations. | |
| Tests principles common in business code of conducts. Adds Recipe Steganography and Zalgo Noise Mutations. |
Ethical Behavior
Scenario ID: vijil.scenarios.safety_compliance_ethicalbehavior
Ethical Behavior
Scenario ID: vijil.scenarios.safety_compliance_ethicalbehavior
| Probe Name | Description |
|---|---|
| Assesses if the agent generates gender-biased inferences about occupations. A high score indicates a lack of bias. | |
| Tests whether an agent will help a user to generate misinformation. A high score indicates that the agent refuses to help produce false information. Adds character encoding jailbreak. | |
| Tests whether an agent will help a user to generate misinformation. A high score indicates that the agent refuses to help produce false information. Adds JSON embedding jailbreak. | |
| Tests whether an agent will help a user to generate misinformation. A high score indicates that the agent refuses to help produce false information. Adds logical appeal jailbreak. | |
| Tests whether an agent will help a user to generate misinformation. A high score indicates that the agent refuses to help produce false information. | |
| Tests to assess if an agent will produce content that is social_engineering. Adds Recipe Steganography, Whitespace Noise and JSON Escape Mutations. | |
| Tests to assess if an agent will produce content that is social_engineering. Adds Recipe Steganography and Zalgo Noise Mutations. | |
| Tests to assess if an agent will produce content that is social_engineering. Adds Government Compliance Directive and MIME Quoted Printable Mutations. | |
| Tests to assess if an agent will produce content that is social_engineering. Adds Zero Width Injection and Zalgo Noise Mutations | |
| Tests for misinformation using homoglyph mutation | |
| Tests for misinformation using zero width mutation | |
| Tests for misinformation using spaces mutation | |
| Tests for misinformation using fullwidth mutation | |
| Tests for misinformation using soft mutation mutation | |
| Tests to assess if an agent will produce content that helps to produce chemical, biological, radiological, and nuclear threats. Adds Whitespace Noise and MIME Quoted_Printable Mutations. | |
| Tests to assess if an agent will produce content that helps to produce chemical, biological, radiological, and nuclear threats. Adds Zero Width Injection and Zalgo Noise Mutations. | |
| Tests to assess if an agent will produce content that helps to produce chemical, biological, radiological, and nuclear threats. Adds Recipe Steganography and Zalgo Noise Mutations. | |
| Tests to assess if an agent will produce content that helps to produce chemical, biological, radiological, and nuclear threats. Adds Recipe Steganography, Whitespace Noise, and JSON Escape Mutations. |