ActuarialPythonchainladder開源Model PersistenceLoss ReservingTesting

What CapeCod.predict Actually Does: chainladder Re-Estimates the Apriori and Returns a Third Number

· 15 min read
Table of Contents
  1. Where Cape Cod's apriori comes from
  2. Three numbers
  3. Taking 0.6894525318 apart
  4. How the maintainer sees it
  5. Freezing the apriori is BF, but only in-sample
  6. Three questions to ask before calling predict
  7. Current status (2026-09-26)
  8. Sources

The short answer: chainladder-python's CapeCod.predict does not apply your saved model to new data. Every call re-estimates the apriori from the new data passed in, but keeps the development pattern from fit time. On the ukmotor sample, one model gives three different aprioris: 0.6572660657 fitted on the prior diagonal, 0.6894525318 returned by predict on the new diagonal, and 0.7062253126 from a clean refit on the new data. The third number is neither the model you saved nor the model you would build today. When we opened the thread we wrote "No bug here". Whether CapeCod should re-estimate is a question the maintainer has handed to the community, and as of September 26, 2026 it is still undecided.

There is only one lesson to take away: before calling any predict, work out which quantities are stored and which are recomputed on every call.

Our role, disclosed first. chainladder-python lives under casact, the official GitHub organization of the Casualty Actuarial Society (CAS). The thread #1274 was opened by us (GitHub account ppcvote), but the first person in the thread to point out that CapeCod re-estimates on every inference was maintainer henrydingliu (September 6, 17:58 UTC). Our comment an hour later quantified it as the three numbers above. All of the ukmotor numbers below were re-run on the current main (commit 08ded050) and on PyPI chainladder==0.10.1, and three other historical commits give results identical to every digit. The hand rebuild, the trend=0.05 comparison and the #1385 example were run on 08ded050.

Where Cape Cod's apriori comes from

In the BornhuetterFerguson method (BF from here on), you supply the expected loss ratio, the apriori. The Cape Cod method does not let you supply it. It estimates it from the data: the sum of reported losses across origin years, divided by the sum of "used-up exposure", which is each origin year's exposure divided by its cumulative development factor (CDF). With trend=0, decay=1 and no on-level adjustment:

apriori = sum(latest diagonal) / sum(exposure / CDF)

The official user guide uses the same formula when it derives the apriori implied by CapeCod (docs/user_guide/methods.ipynb, cell 29).

The formula has three ingredients: the latest diagonal, the exposure and the CDF. Which period each of the three comes from when predict is called is the question this whole post answers.

Three numbers

ukmotor is a sample that ships with chainladder: a 7×7 cumulative triangle by origin year, latest valuation date 2013-12-31. It has no premium column, so the exposure is synthetic: a flat 20000 per origin year, written the same way as the example in the CapeCod.predict docstring.

import chainladder as cl
tr = cl.load_sample("ukmotor")
tr_prior = tr[tr.valuation < tr.valuation_date]           # drop the latest diagonal
exp_prior = cl.Chainladder().fit(tr_prior).ultimate_ * 0 + 20000
exp_now   = cl.Chainladder().fit(tr).ultimate_ * 0 + 20000
cc    = cl.CapeCod().fit(tr_prior, sample_weight=exp_prior)
pred  = cc.predict(tr, sample_weight=exp_now)
refit = cl.CapeCod().fit(tr, sample_weight=exp_now)
Scenario apriori_
A Fit on the triangle as of end-2012 (6 origin years) 0.6572660657
B A's model, predict on the triangle as of end-2013 (7 origin years) 0.6894525318
C Clean refit on the end-2013 triangle 0.7062253126

Our September 6 comment put it like this: "predict returns a third number, neither the fitted apriori nor a clean refit."

CapeCod.predict has a branch that infers the fitted grain, and the one-character bug from the previous post sits there. In this example both sides have index levels ['Total'], the branch condition evaluates to an empty set, and no warning of any kind is raised at run time, so the gap between the three numbers does not come from that path.

Also, the docstring example uses trend=0.05, while the setup above uses the default trend=0. Switch to 0.05 and the three numbers become 0.7605768615, 0.8162389074 and 0.8360961065, still three different values.

Taking 0.6894525318 apart

We rebuilt each one by hand, origin year by origin year, with the formula above, and all three results match chainladder's output to 10 decimal places.

Latest diagonal total Sum of exposure / CDF CDFs used apriori
A fit 58994 89756.6497 fit time 0.6572660657
B predict 75672 109756.6497 fit time 0.6894525318
C refit 75672 107149.9402 new 0.7062253126

Comparing them in pairs:

  • B and C use the same losses (75672) and the same exposure. The only difference is the CDF.
  • A and B use the same CDF. The only difference is the data: one new diagonal, plus origin year 2013.

So 0.6894525318 is not an average of, or a compromise between, the other two. It is the CapeCod formula taking in today's losses and exposure, paired with last period's development pattern. In this example it happens to land between the two, but we only tested two settings, trend 0 and 0.05, so we cannot infer that it always lands in the middle.

The most concrete cell is origin year 2008, which has developed to 72 months as of end-2013. The end-2012 triangle used for the fit never observed development from 72 to 84 months, so the fit-time pattern is 1.000000 for that step, and B uses 1.000000. The refit triangle can see that step, and C uses 1.027530.

The corresponding source is in CapeCod.predict. Here it is verbatim at commit 08ded050 (the if is on line 329; in PyPI 0.10.1 it is line 324):

X_new = X.copy()
_, X_new.ldf_ = self.intersection(X_new, self.ldf_)
# If model was fit at a higher grain, then need to aggregate predicted aprioris too
if len(set(sample_weight.key_labels) - set(self.apriori_.key_labels)) > 0:
    apriori_, detrended_apriori_ = self._get_capecod_aprioris(
        X_new.groupby(self.apriori_.key_labels).sum(),
        sample_weight.groupby(self.apriori_.key_labels).sum(),
    )
else:
    apriori_, detrended_apriori_ = self._get_capecod_aprioris(
        X_new, sample_weight
    )

The intersection line runs first: it attaches the fitted ldf_ to the new data, X_new. Then both branches of the if call _get_capecod_aprioris, recomputing the apriori from this X_new and the sample_weight passed in. No path directly reuses the apriori_ stored at fit time.

That the recomputation uses the fit-time CDFs, we confirmed with two run results. First, the hand calculation for row B above, plugging in the fit-time CDFs, matches chainladder's output to 10 decimal places. Second, pred.ldf_ == cc.ldf_ prints True.

Reusing the fitted ldf_ is not unusual in itself. As we noted in #1274 on September 4, Chainladder and BornhuetterFerguson do the same in predict. What sets CapeCod apart is that it also re-estimates the apriori, so the result mixes inputs from two periods.

How the maintainer sees it

#1274 was not originally about this. Its title is "CapeCod.predict infers the fitted grain from key_labels: should that inference exist?", asking whether grain inference should exist at all, and the first line of the body says "No bug here". By September 6 the discussion had reached the point where henrydingliu (one of the project's CODEOWNERS, and the person who merged our #1275) wrote:

the current implementation of capecod re-estimates apriori_ on every inference. there is essentially no model persistence. as in, i can't save the capdecod from last quarter and reapply it this quarter.

("capdecod" is the original spelling.) In plain terms: the current CapeCod re-estimates apriori_ on every inference, there is essentially no model persistence, and a CapeCod saved last quarter cannot be applied to this quarter.

Later the same day, he added a practical observation and the current state of the package:

in practice, the a priori for a BF method is often determined based on some historical chainladder result.

this package currently doesn't support putting this entire analysis into a pipeline. something we'll have to revisit in the future

In the same comment, he laid out two paths for the community to choose between:

if consensus is 'as is', then we make explicit disclaimer in the docstring around the re-estimation behavior - if consensus is 'no re-estimation', then a bigger refactor effort would be needed.

If the consensus is to keep things as they are, the docstring states the re-estimation explicitly. If the consensus is no re-estimation, a bigger refactor is needed. As for whether to warn on every call, both sides agreed that putting it in the docstring is enough. Our reason at the time: a warning that fires on every call is one people switch off.

On September 17 he opened #1385, titled "[BUG/BRK] Is CapeCod a property or an estimator?", with the core question written like this:

what is the capecod method at its core? is it just a way of calculating a bf apriori? do we expect that apriori to update? if we do expect that apriori to change, capecod ultimate actually becomes a property, rather than an estimated result.

Put another way: what is Cape Cod at its core? Is it only a way to compute a BF apriori? Do we expect the apriori to update? If we do, CapeCod's ultimate becomes a "property" rather than an estimated result.

He attached a small example to #1385. We re-ran it unchanged on 08ded050, and the output matches what the issue prints: the model fits an apriori of 0.625 on the first state, and predict on a new state with only two origin years returns 0.416667. On the same new state, Chainladder's predict gives 3750 and 2500, BF gives 4500 and 4500, and CapeCod gives 3500 and 3166.67. The example is his, not ours.

Freezing the apriori is BF, but only in-sample

Our September 6 comment also contained this line:

Persistence and re-estimation are the same fork, and taking one means opting out of the thing CapeCod does.

Persistence and re-estimation are one fork in the road: pick either side, and you have given up the thing CapeCod does. That is our framing. The maintainer did not quote it and did not say whether he agrees.

The equality behind it is not a new finding. The official user guide already says default CapeCod "can be emulated by" BornhuetterFerguson. In the source, CapeCod.fit builds BF's expectation directly and hands it to BF:

self.expectation_ = sample_weight * self.detrended_apriori_

So if you multiply detrended_apriori_ by the exposure and feed it to BF, on the same data both sides give a total ultimate of 78871.9279, with a cell-by-cell difference of 0.0. This holds by construction, and on September 13 we added ourselves: "That is close to tautological".

What matters is out of sample. Carry the stored detrended_apriori_ to the next diagonal (BF's sample_weight set to the new exposure times the stored value), and the output is:

Item Result
Origin years covered by the stored detrended_apriori_ 2007 to 2012
sample_weight for origin year 2013 nan
ultimate_ for origin year 2013 6283.00, equal to reported losses
ibnr_ for origin year 2013 NaN
Total ultimate, first six origin years frozen 80242.93, CapeCod.predict 80774.45, about 0.66% apart

The newest origin year gets no IBNR at all, and ultimate_ contains no NaN, so ultimate_ alone does not show that anything is missing. This lines up with the maintainer's correction on September 6: "detrended_apriori_ is origin-specific because the trend and olf vectors are origin-specific." The stored values are tied to the old origin years, and a new origin year has no counterpart.

The conditions on the 0.66% must be stated with it: one small public triangle, paired with a flat synthetic exposure of 20000. It illustrates the mechanism and cannot be taken as the size of the error on a real book.

Three questions to ask before calling predict

This example applies to any model that is fitted once and then used for repeated predict calls, actuarial or not:

  1. Which quantities are stored, and which are recomputed on every call? CapeCod stores ldf_ and recomputes apriori_. If the documentation does not make this clear, print the key attributes from fit and from predict side by side and compare.
  2. Which quantity accounts for the difference between predict and a clean refit? Run each once on the same new data and compare item by item. In this example the whole difference comes from the CDF; the losses and the exposure are identical.
  3. Is the thing you want to persist defined for the new data? A parameter indexed by origin year has no value when a new origin year arrives. In this example the consequence is zero IBNR for the newest origin year, with no error raised. A development pattern indexed by line of business has the same gap: when the prediction data carries a line of business the model never saw, Chainladder's predict drops those rows without raising.

Until the community reaches a conclusion on #1385, if your workflow is "apply last quarter's CapeCod to this quarter", at minimum know that the apriori_ predict returns is no longer last quarter's.

Current status (2026-09-26)

  • #1274: open, 14 comments. The last one is the maintainer's, from September 17: "new issue raised at #1385. thanks for all the help. we'll be able to finally close this out soon after some community input."
  • #1385: open, label Triage Pending, zero replies.
  • #1306: a separate PR of ours. It deals with a warning during grain inference, not with re-estimation. It is not merged, is awaiting review, and has conflicts with main.

No direction has been decided.

Sources

  • casact/chainladder-python issue #1274 (opened by us). Comments quoted: maintainer 2026-09-06 17:58 UTC (comment 5561081469), ours 2026-09-06 18:58 UTC (comment 5561426764), maintainer 2026-09-06 20:55 UTC (comment 5562110562), ours 2026-09-13 (comment 5652358664), maintainer 2026-09-17 (comment 5720823903). As checked on 2026-09-26, none of the 14 comments had been edited since posting.
  • Issue #1385 (opened by henrydingliu, 2026-09-17).
  • Official user guide, Methods; corresponding source file docs/user_guide/methods.ipynb, cells 28 and 29.
  • Source code: CapeCod.fit and CapeCod.predict in chainladder/methods/capecod.py, commit 08ded050188728f155f6ad71b7d7b7586da6eea2 (main, 2026-09-25). The if in predict is line 329 at this commit and line 324 in PyPI 0.10.1; the capecod.py:325 cited in the #1274 body is from an earlier commit.
  • Reproduction: dataset cl.load_sample("ukmotor"), trend=0, flat exposure of 20000, code as in the snippet above. Environment: a standalone venv on Python 3.12.4 (py -3.12 -m venv venv), numpy 2.5.3, pandas 3.0.6. A plain pip install chainladder==0.10.1 is enough to reproduce. Main and three other commits were run with PYTHONPATH pointing at that commit's source, using the same code: 91a942f1 (the #1275 merge point), cd24e1d2 (main at the time of the September 6 comment) and 6b51f649 (the commit mentioned in the September 13 comment). Every number is the same across all five versions.
  • Hand rebuild of the mechanism and the #1385 example: same environment, commit 08ded050, run on 2026-09-26.

FAQ

Does CapeCod.predict reuse the apriori computed at fit time?

No. Every predict call recomputes the apriori from the new triangle and new exposure passed in, and reuses only the fit-time development factors, ldf_. In the ukmotor example, the fitted apriori is 0.6572660657 and predict returns 0.6894525318.

Then is the predict result the same as refitting on the new data?

Not that either. A clean refit on the new data gives 0.7062253126. predict and the refit use exactly the same losses (latest diagonal total 75672) and exposure. The only difference is the development pattern: predict uses the prior period's cumulative development factors, the refit uses new ones. For example, for origin year 2008 at 72 months, predict uses 1.000000 and the refit uses 1.027530.

Is this a bug in chainladder?

When we opened #1274, the first line of the body said No bug here, raising it as a design question. On 2026-09-17 maintainer henrydingliu opened a separate issue, #1385 (titled Is CapeCod a property or an estimator?), asking the community to decide whether CapeCod should stay as it is with the docstring stating that it re-estimates, or change to not re-estimating (which would need a bigger refactor). As of 2026-09-26, #1385 has no replies.

Is freezing CapeCod's apriori and reusing it the same as BornhuetterFerguson?

Only in-sample. The official user guide already says default CapeCod can be emulated by BornhuetterFerguson, and on the same data both give a total ultimate of 78871.9279. But carry the stored detrended_apriori_ to the next diagonal and the new origin year 2013 has no corresponding value, so its ultimate equals the reported losses of 6283.00, with no IBNR booked at all.

How big is this gap in practice?

We have no data to answer that. The only thing we measured is ukmotor, a small public triangle, with a flat 20000 exposure copied from the docstring example: the total ultimate for the first six origin years differs by about 0.66%. It illustrates the mechanism and does not represent the reserve error on a real book.

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.