AI 安全
28 articles about "AI 安全".
ChatGPT Privacy Center Is a Map, Not a New Switch: Reading It Through Four Questions
OpenAI announced Privacy Center in the ChatGPT release notes on 21 September 2026. The official help article states that the controls remain in Settings, and that opening it does not change your settings or delete your data. Checked against four questions (default, scope, retention, revocation): the default has not moved, and there is no topic on images or faces; the multi-account connections added four days earlier widen the data sources a single conversation can reach.
Regex Can't Tell "Said It" From "Warned Against It": False Positives, Misses, and Our Overturned Fix in LLM Output Scanning
Matching LLM replies with regex to catch XSS or SQL injection strings also fails replies that name the hazard in order to warn against it: on an unmerged open-source Giskard PR, all 5 such advice sentences tripped the check. We proposed anchoring the patterns to line start and false positives fell to 0. Then the PR author added one ordinary phrasing, a lead-in on the same line as the payload, and the anchor missed 12 of 24 compliant replies. This post covers why the anchor only held on our sample, how to choose between over-reporting and under-reporting, and why our own output scanner has the same limit.
AI Skills Are an Injection Path Too: Scanning 954 Popular Skills, and an Eighth Question Before Buying an AI System
A September 2026 paper scanned 954 popular AI agent skills. Its tool flagged 17.6%, and reviewers confirmed 84 of 100 sampled findings as latent vulnerabilities. Five skills were attacked for real and 13 of 30 attempts succeeded, even when models recognised the risk. Around the same time Microsoft disclosed two 9.9-rated Semantic Kernel vulnerabilities, and both arrive at the same line: the model is not a security boundary. This post unpacks the qualifiers on each number and adds an eighth question to our seven questions to ask before buying an AI one-person-company system.
ChatGPT's Terms Treat Your Face as "Content": Reading OpenAI's Terms on Uploaded Photos
A paragraph-by-paragraph reading of OpenAI's official terms. The Terms of Use file uploaded images under "Input", the same category as typed text; ownership stays in your name while usage rights also go to OpenAI; in the Service Terms passage on likeness, all six verbs have "you" as their subject; and the training help page explicitly separates services for individuals from business products. Nowhere in the terms does the face get a section of its own. Every quotation was checked against the official pages as they read at the time of writing.
Uploading a Selfie to ChatGPT Is Not a Filter: Your Face Is a Credential You Cannot Reset
Social feeds are full of "ask ChatGPT what I would look like as a Japanese gangster": upload a selfie, get a generated image back, then tag someone with "your turn". This is not a filter. A face is a credential that cannot be reset, consumer accounts by default allow content to be used to improve models, turning the setting off is not retroactive, and deleting a chat does not pull it back out of a training batch. With the original text of the terms, the full pipeline, the line between what can go to the cloud and what should never be uploaded, and the two paths in our own products that touch faces.
If Your Product Takes Photo Uploads, You Are Already Building a Biometric System: Four Questions for Every Image Endpoint
For people building AI products: once there is a path where a user uploads a photo and another face is generated from it, the data category is already identity material, whether the feature copy calls it a dress-up or a filter. Every upload endpoint that takes images needs four answers fixed before launch: default, scope, retention and revocation. Includes a full checklist, plus the two paths in our own products that touch faces and the work we have not finished yet.
What Is GPT-6 Astra? Capabilities, Pricing, API Migration, and Five Safety Numbers to Read Before You Hand It an Agent
OpenAI released GPT-6 Astra on September 3, 2026, positioned as the most intelligent and aligned model in the world and the first to reach the Critical cybersecurity level under OpenAI's own Preparedness Framework. This guide covers what it is, how it differs from GPT-5.6 Sol, how API pricing works, which parameters change on migration, and then lays out the five numbers in the system card that matter for AI agents: an 8.5% indirect prompt injection success rate, 99.99% instruction hierarchy, 0% on honeypots, and one number pointing the other way: chain-of-thought monitorability went down. Then how enterprises should test, and how they should not.
The Scanner Had the Bug It Was Looking For: Auditing Someone Else's Invisible-Character List, Then Our Own
Anthropic's commerce-agents blueprint strips invisible characters with a hand-written list. A Unicode category sweep found 34 Cf code points and four invisible non-Cf characters missing from it, enough to keep a fence label or a role word intact through sanitizing. Then we pointed the same probe at our own code and all three of our scanners missed the same set. On why enumerations expire, how to grade a finding honestly, and how binding rules to wording instead of concepts makes you systematically underrate the systems that got it right.
Why Agentic AI Attack Testing Shouldn't Be One Class Per Attack: The Vector / Framing / Scorer Decomposition
In an agent pipeline the same malicious payload can enter as a tool result, a retrieved document, or a sub-agent message. Write one attack class per entry point and your test harness explodes. This is the three-axis model we posted publicly in microsoft/PyRIT: injection vector is data, framing is a transform, the scorer is the verdict, and why an attack is a placement, not a message.
The July 2026 MCP Server Auth Epidemic: 12 Projects, 19 Advisories, and the Reference SDK Itself
In three weeks of July 2026, at least 12 MCP projects, including the official MCP Python SDK itself, shipped 19 security advisories, all pointing at the same thing: authentication and origin validation treated as optional. This isn't a run of bad luck, it's a structural gap. A source-verified inventory, the five recurring classes and their root cause, and the checklist to run before you connect any MCP server.
OWASP Agentic Top 10 (ASI-01 to ASI-10) Explained for 2026: Real Scenarios, Detection and a Fix Checklist for Each Risk
A risk-by-risk breakdown of the OWASP Agentic Top 10 (ASI-01 to ASI-10): what each risk is, what a real scenario looks like, how to scan for it with open-source tools, and a fix checklist you can copy directly. Written for people actually running AI agents.
OWASP LLM Top 10 Explained (Official 2025 Edition, 2026 Status): How It Differs from the Agentic ASI Top 10 and How Developers Should Defend
Which list should you read when you search for "OWASP LLM Top 10 2026"? The latest official version is still the 2025 edition; what is new in 2026 is the Agentic ASI attack surface. This post walks through all 10 risks one by one, compares LLM and Agentic security, and gives developers three layers of defense they can put into practice.
AI Agent Output Guardrails: How to Stop API Key Leaks Before Your Agent Posts
We sell AI guardrails, yet our own agent was posting publicly with no output checks and could have leaked API keys. How to build an egress guardrail, and how to decide what to hard-block and what only gets a soft warning.
From 6 to 21: The Crypto AI Agent Incident Tracker Goes Live ($52M of Documented Loss)
The 6 incidents from an earlier analysis, expanded to 21 today. $52M total documented loss. Structured data, open-source repo, public page. Built in-flight.
Six Crypto AI Agent Heists: What Static Prompt Analysis Catches, What It Doesn't
An honest root-cause analysis of six prompt-injection incidents that drained crypto AI agents, and a measured assessment of what prompt-defense-audit can and cannot catch.
We Audited 7 Official MCP Servers: 6 Got F
Ran prompt-defense-audit against the 7 official servers in modelcontextprotocol/servers: 12-vector check, OWASP LLM Top 10 mapping. Result: 6 servers scored F, 8 defense vectors at 100% gap rate. Cross-referenced from modelcontextprotocol/servers#3537.
Cisco Merged My PR in 51 Minutes: Why Prompt Defense Is the Next SQL Injection
AI agents and chatbots are growing exponentially, foundation models update every three months, but 78% of production prompts have zero defense lines. From one casual scan to Cisco merging in 51 minutes and Microsoft assigning me an issue: the four months between.
OWASP Agentic Top 10: What Every AI Developer Needs to Know in 2026
OWASP released its Top 10 security risks for AI agent applications in 2026. We break down each risk with real data from scanning 1,646 production system prompts.
One Line to Block 92% of Prompt Injection Attacks
Our Discord AI assistant gets attacked every few days. After scanning 1,646 real AI systems, we built a one-liner defense tool.
We Built Lighthouse for AI Agents: One Command, 25-Vector Security Audit
66% of MCP servers have security findings, but nobody runs a security scan before deploying AI agents. We built ultraprobe: zero deps, zero cost, under 1 second. Its prompt defense scanning technique was merged into Cisco AI Defense's MCP Scanner (PR #146).
12 Submissions, 0 Merges: What I Learned Contributing to Open Source AI Security
We submitted contributions to Cisco, Microsoft, OWASP, and 9 other open source projects. All rejected or ignored. Here's how we went from 0/12 to our first merge.
We Defined an AI Security Standard: AASS v1.0, We Don't Sell Security, We Define It
AI Application Security Standard (AASS) is the first open standard covering AI system defense, website AI visibility, and data protection in a single framework. All tools free and open source.
We Scanned 1,646 Real AI System Prompts. Here's What We Found.
We ran our prompt defense scanner against 1,646 leaked production system prompts from ChatGPT, Claude, Grok, Cursor, Perplexity, and 1,300+ custom GPTs. 97.8% have no indirect injection defense. Average score: 36/100.
Prompt Injection Isn't Your Biggest Risk: We Scanned 517 AI System Prompts and Found 11 Undefended Attack Vectors
Everyone talks about Prompt Injection, but it's just 1 of 12 LLM attack vectors. We scanned 517 AI system prompts with UltraProbe and found they defend against only 3.2 of the 12 on average. Here are the other 11 you're ignoring.
78.3% Score F: Prompt Defense Gap Data from 1,646 Real AI System Prompts
We scanned 1,646 system prompts leaked from GPT Store, ChatGPT, Claude, Cursor and others. The average score was 36/100 and 78.3% scored F. This post uses the data to show how serious each OWASP Agentic Top 10 risk is in the real world.
We Open-Sourced Our Prompt Defense Scanner: 200 Lines of Regex That Replace an LLM
Most AI security tools use LLMs to check LLMs. We built a deterministic prompt defense scanner: 12 attack vectors, pure regex, under 1ms, zero cost. Here's why regex beats AI for this job, and how you can use it today.
How We Defend AI Against Comment Attacks: 5-Layer Prompt Defense in Production
When your AI auto-replies to hundreds of comments daily, Prompt Injection isn't theoretical: it's happening every day. This is the 5-layer defense architecture we validated across 27 accounts.
UltraProbe Is Live: The World's First Free AI Security Scanner That Finds Your LLM Vulnerabilities in 5 Seconds
90% of AI systems are vulnerable to Prompt Injection, yet most developers have no idea. Ultra Lab launches the completely free UltraProbe, covering the OWASP LLM Top 10 attack vectors, making AI security testing accessible to everyone, not just enterprises.