Input Validationsilent failuresTestingPythonActuarial開源chainladder

A Validation Check One Line Too Late Doesn't Raise, It Becomes Decoration: Four Silent Failures in chainladder's predict

· 15 min read
Table of Contents
  1. What predict does
  2. Four ways to be wrong, none with an error message
  3. The shape of the fix was set by the maintainer
  4. Lesson one: legitimate and dangerous looseness share a shape, so find the system's own marker
  5. Lesson two: the check has to sit before the line that rewrites its input
  6. Status (2026-09-26)
  7. AI use disclosure
  8. Sources

The short answer: anyone who writes input validation can take away two things. First, a check placed after any code that rewrites its input does not raise; it goes quietly blind in exactly the case it most needs to see, and every other test stays green. Second, legitimate looseness and dangerous looseness often look exactly the same in the data. Shape cannot separate them, so look for a marker the system writes itself, and write down where that marker goes wrong. Both lessons come from the same predict in the actuarial library chainladder-python, which returned normal-looking results for four kinds of mismatched input.

Disclosure: issue #1288 and the fix PR #1310 were both opened by us (GitHub account ppcvote). As of September 26, 2026, the PR is open, CI is all green, and it has no reviews and no merge. The latest release on PyPI, 0.10.1, and that day's main both still have the original behaviour. Following casact's policy, the PR discloses AI use; the original text is at the end of this post.

What predict does

chainladder-python is the loss-reserving library maintained by the Casualty Actuarial Society (casact). fit learns a development pattern ldf_ from one triangle, and predict(X) applies that pattern to a new triangle X to compute ultimate losses, ultimate_. A triangle has an index (company GRNAME, line of business LOB) and columns (paid CumPaidLoss, incurred IncurLoss). Background on grain is in our earlier breakdown of CapeCod.predict. CapeCod's predict also re-estimates the apriori from the X and exposure passed in, keeping only the fit-time ldf_, so its result mixes inputs from two periods.

predict attaches the pattern to X and then hands both to intersection() to align them. When the two sides do not match, it narrows them to what they have in common; when both sides are a single row, its first line, if len(a) == 1 and len(b) == 1: return a, b, lets them straight through untouched. Its job is alignment, not asking why things do not line up.

Four ways to be wrong, none with an error message

On the PR's base commit e4e06956, with the clrd sample that ships with the library (775 rows, each row one company's triangle on one line of business), using the Chainladder method:

Scenario What we did Result
1. X has a group the model never saw Fit on the total of the 5 lines of business other than wkcomp, predict all 775 rows ultimate_ has only 643 rows, wkcomp's 132 rows are gone, 0 warnings
2. Pattern finer than X Fit company by company on 775 rows, predict the totals of the 6 lines of business A 6-row input gets a 775-row ultimate_, with an extra GRNAME index level
3. Both sides a single row State Farm's comauto pattern, predict State Farm's othliab 1,638,150.65, labelled othliab; othliab with its own pattern is 2,640,829.49, so 38% short, 0 warnings
4. Column mismatch Paid pattern from the line-of-business totals, predict incurred data 228,088,946, with the column named IncurLoss; incurred with its own pattern is 150,105,776, overstated by 52%, 0 warnings

The second one emitted 10 numpy runtime warnings (square root and overflow), none of which mention the grain mismatch. Given the same input as the first, BornhuetterFerguson and CapeCod throw operands could not be broadcast together with shapes (775,1,10,1) (643,1,10,1): there is an error, but it comes from numpy's shape mismatch, and the message names no index value.

The same script run on main 08ded050 and on 0.10.1 from PyPI gives the same output as the base. Shortest reproductions of the first and fourth:

import chainladder as cl
clrd = cl.load_sample("clrd")
tri = clrd["CumPaidLoss"]

train = tri[tri.index["LOB"] != "wkcomp"].groupby("LOB").sum()
pred = cl.Chainladder().fit(cl.Development().fit_transform(train)).predict(tri)
tri.shape[0], pred.ultimate_.shape[0]            # (775, 643)

paid = clrd["CumPaidLoss"].groupby("LOB").sum()
incurred = clrd["IncurLoss"].groupby("LOB").sum()
model = cl.Chainladder().fit(cl.Development().fit_transform(paid))
model.predict(incurred).ultimate_.sum().sum()    # 228,088,946

We have no evidence that anyone's reserves were miscalculated because of this. What these numbers show is that when the input is wrong, the return value still looks like a normal report.

The shape of the fix was set by the maintainer

We opened #1288 on September 4. Maintainer henrydingliu replied on September 5:

i think we can simply error out on the prediction if intersection doesn't return the starting index. i believe this is the most consistent with other sklearn estimators when a feature introduces a new level at inference.

Our first proposal was to put the check inside intersection(). The maintainer mentioned that he had considered promoting intersection to a triangle method (#1037) and was disinclined to add error handling inside it; instead he proposed factoring a validate_ldf out of predict, alongside the existing validate_X and validate_weight. We replied "validate_ldf is better than what I suggested and I withdraw the intersection() version." The PR was written to that shape.

At PR head bc2563dd, the first three now raise, for example X has index values the model was not fit on: ['wkcomp'] and The fitted pattern has index levels that X does not: ['GRNAME']. It cannot be applied to a triangle that does not carry them.. The fourth does not, for reasons covered in the next section.

Lesson one: legitimate and dangerous looseness share a shape, so find the system's own marker

The most direct rule is "every index value in X must appear in the model". That rule would break the existing test test_different_backends: it first narrows the data to wkcomp, fits on the single row clrd["CumPaidLoss"].sum(), then predicts wkcomp's 132 companies. Applying a group's total pattern to each member of that group is legitimate use.

The dangerous kind: the comauto line-of-business total's pattern, applied to the othliab line-of-business total (the PR's test rejects_a_different_single_group).

In the data the two have the same shape: the pattern is a single row in both, and its index values are not in X. Neither row counts nor index-value comparison can tell them apart.

What tells them apart is a marker chainladder writes itself. When aggregating along axis 0 (sum, mean, median, max, min, prod, var) collapses more than one row into one, the library stamps the index values as "(All)". The first rule in validate_ldf recognises that marker:

if len(ldf) == 1 and set(ldf.index.values.flatten()) == {"(All)"}:
    return

The marker has limits, and the PR description states them plainly: sum() stamps "(All)" on any subset, so "the total of a single line of business" is exempt just as "the total of the whole book" is. Reproduced at PR head, the wkcomp sum pattern applied to comauto returns 157 rows with no error. Telling the two apart would require recording aggregation provenance on the triangle, which the PR description calls "a larger change than this and your call rather than mine".

On the column side, the maintainer drew the line starting from a different system default. A triangle constructed without an index gets the default index "Total", and on September 23 he wrote:

an ibnr predictor trained on an unindexed triangle will work on every other unindex triangle, even if we start to enforce strict index matching.

He extended that precedent to columns and laid out a five-row table, in which row B is "single column | single column | works | precedent from A + ML best practice", while row E, a multi-column fit predicting a column outside it, raises. We changed the PR to match the table (commit f647d352, which limits the column check to fits wider than one column). The result is that at PR head the fourth case still returns 228,088,946, with the column named IncurLoss and no error. Which kind of looseness counts as legitimate is the maintainer's design decision; in the table in our reply we put a number on the cost of this cell: paid -> incurred, raised before, works after, IncurLoss at 228,088,946.

The reusable practice: when deciding whether to let loose input through, first check whether the system itself leaves a mark on the legitimate kind (a sentinel value, a default set at construction). Decide by the mark, not by the shape. Then write down where the mark goes wrong, in the documentation and in code comments.

Lesson two: the check has to sit before the line that rewrites its input

The front half of predict at PR head:

X_new = X.val_to_dev()
# Before the line below, which borrows self.X_'s index when both sides
# are a single row and so would erase what the caller actually passed.
self.validate_ldf(X_new, self.ldf_)
X_new = X_new + (self.X_.val_to_dev().iloc[0, 0].sum(2) * 0)
self.validate_weight(X_new, sample_weight)
X_new.ldf_ = self.ldf_
X_new, X_new.ldf_ = self.intersection(X_new, X_new.ldf_)

(A few lines unrelated to this topic are omitted.) The key is that addition of something multiplied by 0. It goes through _prep_index, which contains this:

if x.kdims.shape[0] == y.kdims.shape[0] == 1 and x.key_labels != y.key_labels:
    kdims = x.kdims if len(x.key_labels) > len(y.key_labels) else y.kdims
    ...
    x.kdims = y.kdims = kdims

When both sides are a single row and their index levels differ, both sides switch to the index of whichever side has more index levels, or the model's side on a tie. After the addition, the identity the caller passed in may already have been replaced by the model's. Measured: with the model fit on the total of the company Aegis Grp (index level GRNAME) and X being the othliab line-of-business total (index level LOB), X's index is othliab before the addition and Aegis Grp after it. On the base, predict returns 3,625,080.92, labelled Aegis Grp.

We got this rule wrong ourselves in the PR. Both the PR description and our comment on the PR on September 20 said: if the two sides share at least one index level name, the result carries the predicted side's label, and only when they share none does it borrow the model's. The reproduction does not support that: Aegis Grp's comauto pattern (index levels GRNAME, LOB) applied to the othliab line-of-business total (index level LOB) shares LOB, and on the base the result is still 3,309,931.25, labelled [['Aegis Grp', 'comauto']]. _prep_index compares the number of index levels, not whether they overlap. On September 26 we corrected the PR description and left a comment saying the original rule was wrong.

We used a script to build two misplaced versions, one with the check on the line after the addition and one with it on the line before intersection. The results are the same:

  • Aegis Grp's comauto pattern (the latest diagonal of this company's comauto sums to 21) applied to the othliab line-of-business total: returns 3,309,931.25, labelled [['Aegis Grp', 'comauto']], no error. othliab with its own pattern is 4,862,567.42, so this result is about 32% lower.
  • The Aegis Grp company total applied to the othliab line-of-business total: 3,625,080.92, labelled Aegis Grp, no error.
  • The State Farm pair (same index levels on both sides) still raises.

The misplaced check is not dead. It still catches the other cases and goes blind only when "both sides are one row and their index levels differ", which is exactly the case it is there to guard.

Can the tests tell? On the PR's first public version, fcbdf4b6, move the check down one line: of the six new tests, the other five stay green and only test_predict_checks_before_the_index_is_borrowed goes red (one per backend, 2 failed, 10 passed). The inputs to those five either have more than one row on at least one side, or have one row on both sides with the same index levels (rejects_a_different_single_group uses the comauto and othliab line-of-business totals, both with index level LOB), so they never enter the index-borrowing branch. At PR head, both misplaced versions give 2 failed, 15 passed, again with only the ordering test red; removing the check entirely is what gives 10 failed.

Every public version of the PR already had the check before the addition, and there is no public record of a misplaced version. The only record is the PR's AI disclosure: the ordering problem was found while "running the suite against earlier attempts that were wrong in each of those ways". The numbers above are our reproduction of "what happens when the check is misplaced", not a reproduction of that earlier version.

The reusable practice:

  1. Between the point where input comes in and the check, list every line that could rewrite the thing you are checking: broadcasting, alignment, filling defaults, type conversion, borrowing an index. Put the check before all of them.
  2. Write one test that deliberately walks into that rewriting branch, then move the check down one line and confirm the test goes red. If everything stays green after the move, your test suite cannot see ordering.
  3. Above the check, write down why it has to be there. When the next person wants to move it somewhere that "looks more natural", the comment and that test stand in the way.

Status (2026-09-26)

  • PR #1310: opened September 7, review requested from three people, 0 reviews. CI on head bc2563dd all passes (ruff, pyright, Python 3.11 to 3.14, pandas3, doctest, codecov, readthedocs), +233/-0, 3 files.
  • On September 24 we reported on the PR that the full suite gives 1274 passed, 7 skipped, against 1260 and 7 on main; the extra 14 are 7 new tests times two backends. Merging the PR locally onto main 08ded050, test_predict.py gives 32 passed, and the reproduction output matches PR head.
  • When the fit has several columns and only one of them is predicted, row D of the table calls for narrowing, but PR head still returns the width from fit time. In our reply we said we would open a separate PR after #1310 merges, and the maintainer answered "narrows as a new PR sounds good"; as of today it has not been opened.
  • Until a fix is released, chainladder users can check for themselves after predict: whether ultimate_ has the same number of rows and the same index as the input, and whether model.X_.columns is the column you think it is.

AI use disclosure

Following casact's AI usage policy, the PR description says:

I used Claude Code on this. It ran the reproductions across the estimators and all three directions, and found both the single-row case and the call-site ordering by running the suite against earlier attempts that were wrong in each of those ways. I reviewed the diff and the tests myself, and the suite was run locally on my machine.

Every number in this post was rerun in an isolated virtual environment; the commands are listed below.

Sources

  • casact/chainladder-python issue #1288, opened 2026-09-04, open: https://github.com/casact/chainladder-python/issues/1288
  • casact/chainladder-python PR #1310, opened 2026-09-07, open, head bc2563dd: https://github.com/casact/chainladder-python/pull/1310 (the maintainer's five-row table from September 23 and our reply are both here; the rule for "which side's label survives" in the PR description and in our 2026-09-20 comment does not match the reproduction; corrected in the description on 2026-09-26 with a comment: https://github.com/casact/chainladder-python/pull/1310#issuecomment-5847411009, see lesson two)
  • Issue #1037, the proposal to promote intersection to a triangle method, open
  • Code (commits e4e06956 and bc2563dd): MethodBase.predict, intersection and validate_ldf in chainladder/methods/base.py; _prep_index in chainladder/core/dunders.py; "(All)" in chainladder/core/pandas.py; the default "Total" in chainladder/core/triangle.py; test_different_backends in chainladder/methods/tests/test_benktander.py
  • Reproduction environment: Python 3.11.6, pandas 3.0.6, numpy 2.4.6; dataset cl.load_sample("clrd"); versions: PR base e4e06956, PR head bc2563dd, first public head fcbdf4b6, main 08ded050, PyPI chainladder 0.10.1
  • Commands: the four scenarios and the (All) exemption use PYTHONPATH=<checkout> python repro.py (the first and fourth are the code above); the tests for the misplaced versions use python -m pytest chainladder/methods/tests/test_predict.py -q -k "predict_rejects or predict_still or predict_checks or predict_columns or misaligned_index"
  • All dates are UTC as shown by GitHub

FAQ

In which cases does chainladder's predict give a wrong result without raising?

On the PR's base commit e4e06956, using the clrd sample, we reproduced four: when the prediction data contains a line of business the model never saw, 775 rows go in and only 643 come back; when the pattern is finer than the data, a 6-row input gets a 775-row ultimate_; when both sides are a single row, State Farm's comauto pattern applied to othliab returns 1,638,150.65, 38% less than the 2,640,829.49 othliab gets with its own pattern; and a paid pattern applied to incurred data returns 228,088,946, overstating the correct 150,105,776 by 52%. None of them produced an error or warning that mentions the mismatch.

Why not simply require every index value in the prediction data to appear in the model?

That rule would reject legitimate use. The existing test test_different_backends fits on the wkcomp sum and then predicts wkcomp's 132 companies, which is applying a group's pattern to the group's members. In the data it has the same shape as 'applying the comauto line's pattern to the othliab line': in both, the pattern is a single row and its index values are not in the prediction data. What tells them apart is the (All) marker chainladder stamps on its own when it aggregates.

What are the limits of using the (All) marker as the criterion?

It is a string comparison, and sum() stamps (All) on any subset. So the total of a single line of business is exempt just like the total of the whole book: at PR head, the wkcomp sum pattern applied to comauto returns 157 rows with no error. The PR description states this limitation plainly; telling the two apart would require recording aggregation provenance on the triangle.

Why does the check stop working when it is placed one line off?

predict contains an addition of something multiplied by 0, and it goes through _prep_index. When both sides are a single row and their index levels differ, it makes both sides share the index of whichever side has more index levels, or the model's on a tie. A check placed after this line sees two triangles that already agree. Measured: Aegis Grp's comauto pattern applied to the othliab line-of-business total returns 3,309,931.25, labelled Aegis Grp and comauto, with no error. The other five new tests stay green; only the one written specifically to pin the order goes red.

Can this fix be used yet?

Not yet. As of September 26, 2026, PR #1310 is open, CI is all green, it has no reviews and it has not been merged; 0.10.1 on PyPI and that day's main 08ded050 both still behave silently. And by the maintainer's design, a single column applied to a single column, paid applied to incurred, will still be allowed after the PR merges.

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.