Prompt Injection
20 articles about "Prompt Injection".
AI Skills Are an Injection Path Too: Scanning 954 Popular Skills, and an Eighth Question Before Buying an AI System
A September 2026 paper scanned 954 popular AI agent skills. Its tool flagged 17.6%, and reviewers confirmed 84 of 100 sampled findings as latent vulnerabilities. Five skills were attacked for real and 13 of 30 attempts succeeded, even when models recognised the risk. Around the same time Microsoft disclosed two 9.9-rated Semantic Kernel vulnerabilities, and both arrive at the same line: the model is not a security boundary. This post unpacks the qualifiers on each number and adds an eighth question to our seven questions to ask before buying an AI one-person-company system.
What Is GPT-6 Astra? Capabilities, Pricing, API Migration, and Five Safety Numbers to Read Before You Hand It an Agent
OpenAI released GPT-6 Astra on September 3, 2026, positioned as the most intelligent and aligned model in the world and the first to reach the Critical cybersecurity level under OpenAI's own Preparedness Framework. This guide covers what it is, how it differs from GPT-5.6 Sol, how API pricing works, which parameters change on migration, and then lays out the five numbers in the system card that matter for AI agents: an 8.5% indirect prompt injection success rate, 99.99% instruction hierarchy, 0% on honeypots, and one number pointing the other way: chain-of-thought monitorability went down. Then how enterprises should test, and how they should not.
The Scanner Had the Bug It Was Looking For: Auditing Someone Else's Invisible-Character List, Then Our Own
Anthropic's commerce-agents blueprint strips invisible characters with a hand-written list. A Unicode category sweep found 34 Cf code points and four invisible non-Cf characters missing from it, enough to keep a fence label or a role word intact through sanitizing. Then we pointed the same probe at our own code and all three of our scanners missed the same set. On why enumerations expire, how to grade a finding honestly, and how binding rules to wording instead of concepts makes you systematically underrate the systems that got it right.
Why Agentic AI Attack Testing Shouldn't Be One Class Per Attack: The Vector / Framing / Scorer Decomposition
In an agent pipeline the same malicious payload can enter as a tool result, a retrieved document, or a sub-agent message. Write one attack class per entry point and your test harness explodes. This is the three-axis model we posted publicly in microsoft/PyRIT: injection vector is data, framing is a transform, the scorer is the verdict, and why an attack is a placement, not a message.
OWASP Agentic Top 10 (ASI-01 to ASI-10) Explained for 2026: Real Scenarios, Detection and a Fix Checklist for Each Risk
A risk-by-risk breakdown of the OWASP Agentic Top 10 (ASI-01 to ASI-10): what each risk is, what a real scenario looks like, how to scan for it with open-source tools, and a fix checklist you can copy directly. Written for people actually running AI agents.
OWASP LLM Top 10 Explained (Official 2025 Edition, 2026 Status): How It Differs from the Agentic ASI Top 10 and How Developers Should Defend
Which list should you read when you search for "OWASP LLM Top 10 2026"? The latest official version is still the 2025 edition; what is new in 2026 is the Agentic ASI attack surface. This post walks through all 10 risks one by one, compares LLM and Agentic security, and gives developers three layers of defense they can put into practice.
From 6 to 21: The Crypto AI Agent Incident Tracker Goes Live ($52M of Documented Loss)
The 6 incidents from an earlier analysis, expanded to 21 today. $52M total documented loss. Structured data, open-source repo, public page. Built in-flight.
Six Crypto AI Agent Heists: What Static Prompt Analysis Catches, What It Doesn't
An honest root-cause analysis of six prompt-injection incidents that drained crypto AI agents, and a measured assessment of what prompt-defense-audit can and cannot catch.
We Audited 7 Official MCP Servers: 6 Got F
Ran prompt-defense-audit against the 7 official servers in modelcontextprotocol/servers: 12-vector check, OWASP LLM Top 10 mapping. Result: 6 servers scored F, 8 defense vectors at 100% gap rate. Cross-referenced from modelcontextprotocol/servers#3537.
Cisco Merged My PR in 51 Minutes: Why Prompt Defense Is the Next SQL Injection
AI agents and chatbots are growing exponentially, foundation models update every three months, but 78% of production prompts have zero defense lines. From one casual scan to Cisco merging in 51 minutes and Microsoft assigning me an issue: the four months between.
OWASP Agentic Top 10: What Every AI Developer Needs to Know in 2026
OWASP released its Top 10 security risks for AI agent applications in 2026. We break down each risk with real data from scanning 1,646 production system prompts.
One Line to Block 92% of Prompt Injection Attacks
Our Discord AI assistant gets attacked every few days. After scanning 1,646 real AI systems, we built a one-liner defense tool.
We Built Lighthouse for AI Agents: One Command, 25-Vector Security Audit
66% of MCP servers have security findings, but nobody runs a security scan before deploying AI agents. We built ultraprobe: zero deps, zero cost, under 1 second. Its prompt defense scanning technique was merged into Cisco AI Defense's MCP Scanner (PR #146).
12 Submissions, 0 Merges: What I Learned Contributing to Open Source AI Security
We submitted contributions to Cisco, Microsoft, OWASP, and 9 other open source projects. All rejected or ignored. Here's how we went from 0/12 to our first merge.
We Scanned 1,646 Real AI System Prompts. Here's What We Found.
We ran our prompt defense scanner against 1,646 leaked production system prompts from ChatGPT, Claude, Grok, Cursor, Perplexity, and 1,300+ custom GPTs. 97.8% have no indirect injection defense. Average score: 36/100.
Prompt Injection Isn't Your Biggest Risk: We Scanned 517 AI System Prompts and Found 11 Undefended Attack Vectors
Everyone talks about Prompt Injection, but it's just 1 of 12 LLM attack vectors. We scanned 517 AI system prompts with UltraProbe and found they defend against only 3.2 of the 12 on average. Here are the other 11 you're ignoring.
78.3% Score F: Prompt Defense Gap Data from 1,646 Real AI System Prompts
We scanned 1,646 system prompts leaked from GPT Store, ChatGPT, Claude, Cursor and others. The average score was 36/100 and 78.3% scored F. This post uses the data to show how serious each OWASP Agentic Top 10 risk is in the real world.
We Open-Sourced Our Prompt Defense Scanner: 200 Lines of Regex That Replace an LLM
Most AI security tools use LLMs to check LLMs. We built a deterministic prompt defense scanner: 12 attack vectors, pure regex, under 1ms, zero cost. Here's why regex beats AI for this job, and how you can use it today.
How We Defend AI Against Comment Attacks: 5-Layer Prompt Defense in Production
When your AI auto-replies to hundreds of comments daily, Prompt Injection isn't theoretical: it's happening every day. This is the 5-layer defense architecture we validated across 27 accounts.
UltraProbe Is Live: The World's First Free AI Security Scanner That Finds Your LLM Vulnerabilities in 5 Seconds
90% of AI systems are vulnerable to Prompt Injection, yet most developers have no idea. Ultra Lab launches the completely free UltraProbe, covering the OWASP LLM Top 10 attack vectors, making AI security testing accessible to everyone, not just enterprises.