OpenAI's GPT-4 Shows Higher Trustworthiness but Vulnerabilities to Jailbreaking and Bias, Research Finds

New research, in partnership with Microsoft, has revealed that OpenAI’s GPT-4 large language model is considered more dependable than its predecessor, GPT-3.5. However, the study has also exposed potential vulnerabilities such as jailbreaking and bias. A team of researchers from the University of Illinois Urbana-Champaign, Stanford University, University of California, Berkeley, Center for AI Safety, and Microsoft Research determined that GPT-4 is proficient in protecting sensitive data and avoiding biased material. Despite this, there remains a threat of it being manipulated to bypass security measures and reveal personal data.

Table of Contents

Trustworthiness Assessment and Vulnerabilities

The researchers conducted a trustworthiness assessment of GPT-4, measuring results in categories such as toxicity, stereotypes, privacy, machine ethics, fairness, and resistance to adversarial tests. GPT-4 received a higher trustworthiness score compared to GPT-3.5. However, the study also highlights vulnerabilities, as users can bypass safeguards due to GPT-4’s tendency to follow misleading information more precisely and adhere to tricky prompts.

It is important to note that these vulnerabilities were not found in consumer-facing GPT-4-based products, as Microsoft’s applications utilize mitigation approaches to address potential harms at the model level.

HONEYSEW Single Double Fold Bias Tape Maker Tool Kit Set, 6MM/9MM/12MM/18MM/25MM Fabric Bias Tape Maker Tools 5 Sizes DIY Sewing Bias Tape Makers for Quilt Binding

DIY Bias Tapes in Minutes-If you are making bias tape for appliqué or any sewing project, this sewing…

As an affiliate, we earn on qualifying purchases.

Testing and Findings

The researchers conducted tests using standard prompts and prompts designed to push GPT-4 to break content policy restrictions without outward bias. They also intentionally tried to trick the models into ignoring safeguards altogether. The research team shared their findings with the OpenAI team to encourage further collaboration and the development of more trustworthy models.

OpenAI's GPT-4 Shows Higher Trustworthiness but Vulnerabilities to Jailbreaking and Bias, Research Finds 5

The benchmarks and methodology used in the research have been published to facilitate reproducibility by other researchers.

Artificial Intelligence and Safety: A Practical Guide for Programmers and Decision Makers

As an affiliate, we earn on qualifying purchases.

Red Teaming and OpenAI’s Response

AI models like GPT-4 often undergo red teaming, where developers test various prompts to identify potential undesirable outcomes. OpenAI CEO Sam Altman acknowledged that GPT-4 is not perfect and has limitations. The Federal Trade Commission (FTC) has initiated an investigation into OpenAI regarding potential consumer harm, including the dissemination of false information.

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications

As an affiliate, we earn on qualifying purchases.

Eldoncard INC Blood Type Test (Complete KIT) – Find Out if You are A, B, O, AB & RH- Results in Minutes – Air Sealed Envelope, Safety Lancet, Micropipette, Cleansing Swab – 1 Pack

Learn your blood type in just a couple of minutes.

As an affiliate, we earn on qualifying purchases.

OpenAI’s GPT-4 Shows Higher Trustworthiness but Vulnerabilities to Jailbreaking and Bias, Research Finds

Up next

Unlocking Small Business Success With Predictive Analytics

Author

James

Trustworthiness Assessment and Vulnerabilities

HONEYSEW Single Double Fold Bias Tape Maker Tool Kit Set, 6MM/9MM/12MM/18MM/25MM Fabric Bias Tape Maker Tools 5 Sizes DIY Sewing Bias Tape Makers for Quilt Binding

Testing and Findings

Artificial Intelligence and Safety: A Practical Guide for Programmers and Decision Makers

Red Teaming and OpenAI’s Response

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications

Eldoncard INC Blood Type Test (Complete KIT) – Find Out if You are A, B, O, AB & RH- Results in Minutes – Air Sealed Envelope, Safety Lancet, Micropipette, Cleansing Swab – 1 Pack

AI Security: The Secret Desire of Every Modern Business

AI in Identity Verification: Safer Logins and Transactions

AI-Powered Malware: When Hackers Use AI Against Us

Must Read! How AI Security Is Making the Internet a Safer Place

11 Best Air Quality Monitors for Healthier Tech Offices in 2026

8 Best Audio Interfaces for AI Music Creators in 2026

11 Best Touchscreen Monitors for AI Workflows in 2026

OpenAI’s GPT-4 Shows Higher Trustworthiness but Vulnerabilities to Jailbreaking and Bias, Research Finds

Up next

Author

James

Trustworthiness Assessment and Vulnerabilities

HONEYSEW Single Double Fold Bias Tape Maker Tool Kit Set, 6MM/9MM/12MM/18MM/25MM Fabric Bias Tape Maker Tools 5 Sizes DIY Sewing Bias Tape Makers for Quilt Binding

Testing and Findings

Artificial Intelligence and Safety: A Practical Guide for Programmers and Decision Makers

Red Teaming and OpenAI’s Response

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications

Eldoncard INC Blood Type Test (Complete KIT) – Find Out if You are A, B, O, AB & RH- Results in Minutes – Air Sealed Envelope, Safety Lancet, Micropipette, Cleansing Swab – 1 Pack

You May Also Like