開源
21 articles about "開源".
What CapeCod.predict Actually Does: chainladder Re-Estimates the Apriori and Returns a Third Number
chainladder-python's CapeCod.predict does not apply the model you saved. On the ukmotor sample, the apriori fitted on the prior diagonal is 0.6572660657, predict on the new diagonal returns 0.6894525318, and a clean refit on the new data gives 0.7062253126. Taking the formula apart, predict uses today's losses and exposure paired with last period's development pattern.
A Validation Check One Line Too Late Doesn't Raise, It Becomes Decoration: Four Silent Failures in chainladder's predict
predict in the actuarial library chainladder-python still returns a result for four kinds of mismatched input: 775 rows go in and 643 come out, 6 rows go in and 775 come out, one row's pattern applied to another row comes out 38% short, and a paid pattern applied to incurred data overstates by 52%, all without an error message. This post unpacks two input-validation lessons from our fix, PR #1310: legitimate and dangerous looseness look identical in the data, so look for a marker the system writes itself; and a check placed after the line that rewrites its input does not raise, it goes blind in exactly the case it most needs to see.
Regex Can't Tell "Said It" From "Warned Against It": False Positives, Misses, and Our Overturned Fix in LLM Output Scanning
Matching LLM replies with regex to catch XSS or SQL injection strings also fails replies that name the hazard in order to warn against it: on an unmerged open-source Giskard PR, all 5 such advice sentences tripped the check. We proposed anchoring the patterns to line start and false positives fell to 0. Then the PR author added one ordinary phrasing, a lead-in on the same line as the payload, and the anchor missed 12 of 24 compliant replies. This post covers why the anchor only held on our sample, how to choose between over-reporting and under-reporting, and why our own output scanner has the same limit.
A Gradient Focus Ring Vanishes in Forced-Colors Mode: Four Lines of CSS in NASA's HDS Core
In CSS forced-colors mode, background-image computes to none for any value that is not url()-based, so a focus ring painted with repeating-linear-gradient layers disappears completely, and if the same rule set outline: none, keyboard users get no indicator at all. That was the bug in NASA's HDS Core, reported by an outside contributor and fixed in PR #185 with four lines of CSS in one mixin. On what forced colors drops and what it recolors, why a mask-image ring survived where the gradient ring did not, and the positive control that kept the measurement honest.
A Promise That Never Settles: How a Failed IndexedDB Write Left the MITRE ATT&CK Search Spinning
A new Promise(async (resolve) => ...) that never took a reject cannot fail. When the IndexedDB chunk write inside it threw, the promise handed to the caller did not reject, it never settled at all, so the .catch the authors had already written could never fire and the search spinner ran until the tab closed. Using mitre-attack/attack-website #637 to explain Promise executor semantics, why clearing ESLint's no-async-promise-executor would not have fixed it, and how to write a test that tells a hang apart from a rejection.
One Character Made 775 Reserve Rows Wrong: Inside chainladder's CapeCod.predict Bug
chainladder-python, maintained by the Casualty Actuarial Society, is the standard library actuaries use for loss reserving. CapeCod.predict() had a condition written as > 1 that should have been > 0, so the most common usage got the wrong apriori: the comauto line fitted 0.569, predict returned 1.252, and all 775 rows were off. An existing test covered the path and still missed it, because its dataset always walked the other branch. On what the Cape Cod method computes, why the line was wrong, how the test missed, and the AI-use disclosure filed under casact's policy.
A Named Parameter Never Reaches **kwargs: Why Every Disaster Photo in HOT's Drone System Was Stored as the Wrong Type
A function declared content_type in its signature and documented it, then called the downstream SDK with positional arguments plus **kwargs. Because content_type is a named parameter, Python bound the value to it and it never reached kwargs, so every post-disaster aerial photo went into S3 as application/octet-stream and browsers downloaded it instead of showing it. Using hotosm/drone-tm #882 to explain Python's binding rules, why metadata in the same function was fine, and a test technique that binds mocks to the real SDK signature.
Seven PRs Merged Into Six Organizations in 24 Days: What the Maintainers Taught Me
Between August 12 and September 5, 2026, 24 days, maintainers at the Casualty Actuarial Society (twice), CERT/CC, the UK AI Security Institute, FINOS, NIST and Humanitarian OpenStreetMap merged seven of my pull requests. This is a line-by-line account: where each bug was, why the existing tests missed it, and what the reviewer corrected. Six unrelated projects, three root causes: Linux-only CI, tests that pass for the wrong reason, and sentinel or binding semantics that betray intuition. Every PR is linked.
The Scanner Had the Bug It Was Looking For: Auditing Someone Else's Invisible-Character List, Then Our Own
Anthropic's commerce-agents blueprint strips invisible characters with a hand-written list. A Unicode category sweep found 34 Cf code points and four invisible non-Cf characters missing from it, enough to keep a fence label or a role word intact through sanitizing. Then we pointed the same probe at our own code and all three of our scanners missed the same set. On why enumerations expire, how to grade a finding honestly, and how binding rules to wording instead of concepts makes you systematically underrate the systems that got it right.
Controlling Claude Code from Telegram: How I Ran It for 60 Days with Zero Downtime on Windows
Turn Telegram into the main console for Claude Code and send it commands from your phone wherever you are. A breakdown of the 4 defense layers and the real incidents behind 60 days of zero downtime on Windows, plus an MIT-licensed open-source toolkit.
We Audited 7 Official MCP Servers: 6 Got F
Ran prompt-defense-audit against the 7 official servers in modelcontextprotocol/servers: 12-vector check, OWASP LLM Top 10 mapping. Result: 6 servers scored F, 8 defense vectors at 100% gap rate. Cross-referenced from modelcontextprotocol/servers#3537.
Cisco Merged My PR in 51 Minutes: Why Prompt Defense Is the Next SQL Injection
AI agents and chatbots are growing exponentially, foundation models update every three months, but 78% of production prompts have zero defense lines. From one casual scan to Cisco merging in 51 minutes and Microsoft assigning me an issue: the four months between.
One Line to Block 92% of Prompt Injection Attacks
Our Discord AI assistant gets attacked every few days. After scanning 1,646 real AI systems, we built a one-liner defense tool.
We Built Lighthouse for AI Agents: One Command, 25-Vector Security Audit
66% of MCP servers have security findings, but nobody runs a security scan before deploying AI agents. We built ultraprobe: zero deps, zero cost, under 1 second. Its prompt defense scanning technique was merged into Cisco AI Defense's MCP Scanner (PR #146).
12 Submissions, 0 Merges: What I Learned Contributing to Open Source AI Security
We submitted contributions to Cisco, Microsoft, OWASP, and 9 other open source projects. All rejected or ignored. Here's how we went from 0/12 to our first merge.
From Zero to Contributing Code to Microsoft: A Non-Engineer's 4-Month Journey
4 months ago I couldn't write a single line of code. Now my PR has been merged into Microsoft's AI governance toolkit. This isn't a genius story. It's a path anyone can follow in the AI era.
We Defined an AI Security Standard: AASS v1.0, We Don't Sell Security, We Define It
AI Application Security Standard (AASS) is the first open standard covering AI system defense, website AI visibility, and data protection in a single framework. All tools free and open source.
We Scanned 1,646 Real AI System Prompts. Here's What We Found.
We ran our prompt defense scanner against 1,646 leaked production system prompts from ChatGPT, Claude, Grok, Cursor, Perplexity, and 1,300+ custom GPTs. 97.8% have no indirect injection defense. Average score: 36/100.
Discord Community From 0 to 146 Members: A Solo Founder's Playbook (With 3 AI Bots)
How does one person build a 146-member Discord community in 10 days? Answer: 3 AI bots + 1 welcome system + $0 ad budget. This is the full SOP from creating the server to retaining members.
78.3% Score F: Prompt Defense Gap Data from 1,646 Real AI System Prompts
We scanned 1,646 system prompts leaked from GPT Store, ChatGPT, Claude, Cursor and others. The average score was 36/100 and 78.3% scored F. This post uses the data to show how serious each OWASP Agentic Top 10 risk is in the real world.
We Open-Sourced Our Prompt Defense Scanner: 200 Lines of Regex That Replace an LLM
Most AI security tools use LLMs to check LLMs. We built a deterministic prompt defense scanner: 12 attack vectors, pure regex, under 1ms, zero cost. Here's why regex beats AI for this job, and how you can use it today.