New research, in partnership with Microsoft, has revealed that OpenAI’s GPT-4 large language model is considered more dependable than its predecessor, GPT-3.5. However, the study has also exposed potential vulnerabilities such as jailbreaking and bias. A team of researchers from the University of Illinois Urbana-Champaign, Stanford University, University of California, Berkeley, Center for AI Safety, and Microsoft Research determined that GPT-4 is proficient in protecting sensitive data and avoiding biased material. Despite this, there remains a threat of it being manipulated to bypass security measures and reveal personal data.

OpenAIs GPT-4 Shows Higher Trustworthiness but Vulnerabilities to Jailbreaking and Bias, Research Finds

Trustworthiness Assessment and Vulnerabilities

The researchers conducted a trustworthiness assessment of GPT-4, measuring results in categories such as toxicity, stereotypes, privacy, machine ethics, fairness, and resistance to adversarial tests. GPT-4 received a higher trustworthiness score compared to GPT-3.5. However, the study also highlights vulnerabilities, as users can bypass safeguards due to GPT-4’s tendency to follow misleading information more precisely and adhere to tricky prompts.

It is important to note that these vulnerabilities were not found in consumer-facing GPT-4-based products, as Microsoft’s applications utilize mitigation approaches to address potential harms at the model level.

HONEYSEW Single Double Fold Bias Tape Maker Tool Kit Set, 6MM/9MM/12MM/18MM/25MM Fabric Bias Tape Maker Tools 5 Sizes DIY Sewing Bias Tape Makers for Quilt Binding

HONEYSEW Single Double Fold Bias Tape Maker Tool Kit Set, 6MM/9MM/12MM/18MM/25MM Fabric Bias Tape Maker Tools 5 Sizes DIY Sewing Bias Tape Makers for Quilt Binding

  • DIY Bias Tape in Minutes: Create custom bias tapes quickly for sewing projects
  • Complete Bias Tape Set: Includes 5 tape makers, binding foot, awl, pins
  • Adjustable Bias Tape Binding Foot: Fits all low shank domestic sewing machines

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing and Findings

The researchers conducted tests using standard prompts and prompts designed to push GPT-4 to break content policy restrictions without outward bias. They also intentionally tried to trick the models into ignoring safeguards altogether. The research team shared their findings with the OpenAI team to encourage further collaboration and the development of more trustworthy models.

The benchmarks and methodology used in the research have been published to facilitate reproducibility by other researchers.

LLM Security in Practice: Essential AI Safety Practices and Attack Prevention (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)

LLM Security in Practice: Essential AI Safety Practices and Attack Prevention (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Red Teaming and OpenAI’s Response

AI models like GPT-4 often undergo red teaming, where developers test various prompts to identify potential undesirable outcomes. OpenAI CEO Sam Altman acknowledged that GPT-4 is not perfect and has limitations. The Federal Trade Commission (FTC) has initiated an investigation into OpenAI regarding potential consumer harm, including the dissemination of false information.

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Artificial Intelligence in Practice Professional Documentation Kit: A 4-Part Practical AI Toolkit for Safe Learning, Responsible Use, Prompts, Templates, and Future-Ready Skills

Artificial Intelligence in Practice Professional Documentation Kit: A 4-Part Practical AI Toolkit for Safe Learning, Responsible Use, Prompts, Templates, and Future-Ready Skills

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Prompt Injection Is a Bigger Deal Than It Sounds

On the surface, prompt injection seems simple, but its true risks could secretly undermine AI security and trust—exploring reveals how vulnerable AI really is.

Zero Trust With AI: Smarter Access Control Systems

Inevitably, AI-driven Zero Trust systems are transforming access control—discover how these innovations can redefine your security strategy.

Boosting RetAIl Success With AI Technology Adoption

We’ve discovered the secret to achieving quick success in retail: integrating AI…

Safeguarding Your Personal Data: Expert Tips for AI Privacy

As technology progresses, safeguarding our personal data is becoming more and more…