New research, in partnership with Microsoft, has revealed that OpenAI’s GPT-4 large language model is considered more dependable than its predecessor, GPT-3.5. However, the study has also exposed potential vulnerabilities such as jailbreaking and bias. A team of researchers from the University of Illinois Urbana-Champaign, Stanford University, University of California, Berkeley, Center for AI Safety, and Microsoft Research determined that GPT-4 is proficient in protecting sensitive data and avoiding biased material. Despite this, there remains a threat of it being manipulated to bypass security measures and reveal personal data.

OpenAIs GPT-4 Shows Higher Trustworthiness but Vulnerabilities to Jailbreaking and Bias, Research Finds

Trustworthiness Assessment and Vulnerabilities

The researchers conducted a trustworthiness assessment of GPT-4, measuring results in categories such as toxicity, stereotypes, privacy, machine ethics, fairness, and resistance to adversarial tests. GPT-4 received a higher trustworthiness score compared to GPT-3.5. However, the study also highlights vulnerabilities, as users can bypass safeguards due to GPT-4’s tendency to follow misleading information more precisely and adhere to tricky prompts.

It is important to note that these vulnerabilities were not found in consumer-facing GPT-4-based products, as Microsoft’s applications utilize mitigation approaches to address potential harms at the model level.

HONEYSEW Single Double Fold Bias Tape Maker Tool Kit Set, 6MM/9MM/12MM/18MM/25MM Fabric Bias Tape Maker Tools 5 Sizes DIY Sewing Bias Tape Makers for Quilt Binding

HONEYSEW Single Double Fold Bias Tape Maker Tool Kit Set, 6MM/9MM/12MM/18MM/25MM Fabric Bias Tape Maker Tools 5 Sizes DIY Sewing Bias Tape Makers for Quilt Binding

DIY Bias Tapes in Minutes-If you are making bias tape for appliqué or any sewing project, this sewing…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing and Findings

The researchers conducted tests using standard prompts and prompts designed to push GPT-4 to break content policy restrictions without outward bias. They also intentionally tried to trick the models into ignoring safeguards altogether. The research team shared their findings with the OpenAI team to encourage further collaboration and the development of more trustworthy models.

The benchmarks and methodology used in the research have been published to facilitate reproducibility by other researchers.

Artificial Intelligence Safety and Security (Chapman & Hall/CRC Artificial Intelligence and Robotics Series)

Artificial Intelligence Safety and Security (Chapman & Hall/CRC Artificial Intelligence and Robotics Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Red Teaming and OpenAI’s Response

AI models like GPT-4 often undergo red teaming, where developers test various prompts to identify potential undesirable outcomes. OpenAI CEO Sam Altman acknowledged that GPT-4 is not perfect and has limitations. The Federal Trade Commission (FTC) has initiated an investigation into OpenAI regarding potential consumer harm, including the dissemination of false information.

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Eldoncard INC Blood Type Test (Complete KIT) - Find Out if You are A, B, O, AB & RH- Results in Minutes - Air Sealed Envelope, Safety Lancet, Micropipette, Cleansing Swab - 1 Pack

Eldoncard INC Blood Type Test (Complete KIT) – Find Out if You are A, B, O, AB & RH- Results in Minutes – Air Sealed Envelope, Safety Lancet, Micropipette, Cleansing Swab – 1 Pack

Learn your blood type in just a couple of minutes.

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Adversarial AI: When Attackers Trick the Algorithms

Meta description: “Many adversarial AI attacks subtly deceive algorithms, and uncovering their evolving tactics is essential to securing AI systems.

Ensure AI Algorithms RemAIn Reliable With These Strategies

Are you fed up with AI algorithms that are not dependable? We…

AI Security: The Untold Story of Your Data’s Best Defender

I’m here to reveal the hidden tale of your data’s ultimate guardian:…

Is Your Data Safe? The Emotional Impact of AI Security

As a cybersecurity expert using artificial intelligence, I can confidently say that…