What CapeCod.predict Actually Does: chainladder Re-Estimates the Apriori and Returns a Third Number
Table of Contents
The short answer: chainladder-python's CapeCod.predict does not apply your saved model to new data. Every call re-estimates the apriori from the new data passed in, but keeps the development pattern from fit time. On the ukmotor sample, one model gives three different aprioris: 0.6572660657 fitted on the prior diagonal, 0.6894525318 returned by predict on the new diagonal, and 0.7062253126 from a clean refit on the new data. The third number is neither the model you saved nor the model you would build today. When we opened the thread we wrote "No bug here". Whether CapeCod should re-estimate is a question the maintainer has handed to the community, and as of September 26, 2026 it is still undecided.
There is only one lesson to take away: before calling any predict, work out which quantities are stored and which are recomputed on every call.
Our role, disclosed first. chainladder-python lives under casact, the official GitHub organization of the Casualty Actuarial Society (CAS). The thread #1274 was opened by us (GitHub account ppcvote), but the first person in the thread to point out that CapeCod re-estimates on every inference was maintainer henrydingliu (September 6, 17:58 UTC). Our comment an hour later quantified it as the three numbers above. All of the ukmotor numbers below were re-run on the current main (commit 08ded050) and on PyPI chainladder==0.10.1, and three other historical commits give results identical to every digit. The hand rebuild, the trend=0.05 comparison and the #1385 example were run on 08ded050.
Where Cape Cod's apriori comes from
In the BornhuetterFerguson method (BF from here on), you supply the expected loss ratio, the apriori. The Cape Cod method does not let you supply it. It estimates it from the data: the sum of reported losses across origin years, divided by the sum of "used-up exposure", which is each origin year's exposure divided by its cumulative development factor (CDF). With trend=0, decay=1 and no on-level adjustment:
apriori = sum(latest diagonal) / sum(exposure / CDF)
The official user guide uses the same formula when it derives the apriori implied by CapeCod (docs/user_guide/methods.ipynb, cell 29).
The formula has three ingredients: the latest diagonal, the exposure and the CDF. Which period each of the three comes from when predict is called is the question this whole post answers.
Three numbers
ukmotor is a sample that ships with chainladder: a 7×7 cumulative triangle by origin year, latest valuation date 2013-12-31. It has no premium column, so the exposure is synthetic: a flat 20000 per origin year, written the same way as the example in the CapeCod.predict docstring.
import chainladder as cl
tr = cl.load_sample("ukmotor")
tr_prior = tr[tr.valuation < tr.valuation_date] # drop the latest diagonal
exp_prior = cl.Chainladder().fit(tr_prior).ultimate_ * 0 + 20000
exp_now = cl.Chainladder().fit(tr).ultimate_ * 0 + 20000
cc = cl.CapeCod().fit(tr_prior, sample_weight=exp_prior)
pred = cc.predict(tr, sample_weight=exp_now)
refit = cl.CapeCod().fit(tr, sample_weight=exp_now)
| Scenario | apriori_ | |
|---|---|---|
| A | Fit on the triangle as of end-2012 (6 origin years) | 0.6572660657 |
| B | A's model, predict on the triangle as of end-2013 (7 origin years) | 0.6894525318 |
| C | Clean refit on the end-2013 triangle | 0.7062253126 |
Our September 6 comment put it like this: "predict returns a third number, neither the fitted apriori nor a clean refit."
CapeCod.predict has a branch that infers the fitted grain, and the one-character bug from the previous post sits there. In this example both sides have index levels ['Total'], the branch condition evaluates to an empty set, and no warning of any kind is raised at run time, so the gap between the three numbers does not come from that path.
Also, the docstring example uses trend=0.05, while the setup above uses the default trend=0. Switch to 0.05 and the three numbers become 0.7605768615, 0.8162389074 and 0.8360961065, still three different values.
Taking 0.6894525318 apart
We rebuilt each one by hand, origin year by origin year, with the formula above, and all three results match chainladder's output to 10 decimal places.
| Latest diagonal total | Sum of exposure / CDF | CDFs used | apriori | |
|---|---|---|---|---|
| A fit | 58994 | 89756.6497 | fit time | 0.6572660657 |
| B predict | 75672 | 109756.6497 | fit time | 0.6894525318 |
| C refit | 75672 | 107149.9402 | new | 0.7062253126 |
Comparing them in pairs:
- B and C use the same losses (75672) and the same exposure. The only difference is the CDF.
- A and B use the same CDF. The only difference is the data: one new diagonal, plus origin year 2013.
So 0.6894525318 is not an average of, or a compromise between, the other two. It is the CapeCod formula taking in today's losses and exposure, paired with last period's development pattern. In this example it happens to land between the two, but we only tested two settings, trend 0 and 0.05, so we cannot infer that it always lands in the middle.
The most concrete cell is origin year 2008, which has developed to 72 months as of end-2013. The end-2012 triangle used for the fit never observed development from 72 to 84 months, so the fit-time pattern is 1.000000 for that step, and B uses 1.000000. The refit triangle can see that step, and C uses 1.027530.
The corresponding source is in CapeCod.predict. Here it is verbatim at commit 08ded050 (the if is on line 329; in PyPI 0.10.1 it is line 324):
X_new = X.copy()
_, X_new.ldf_ = self.intersection(X_new, self.ldf_)
# If model was fit at a higher grain, then need to aggregate predicted aprioris too
if len(set(sample_weight.key_labels) - set(self.apriori_.key_labels)) > 0:
apriori_, detrended_apriori_ = self._get_capecod_aprioris(
X_new.groupby(self.apriori_.key_labels).sum(),
sample_weight.groupby(self.apriori_.key_labels).sum(),
)
else:
apriori_, detrended_apriori_ = self._get_capecod_aprioris(
X_new, sample_weight
)
The intersection line runs first: it attaches the fitted ldf_ to the new data, X_new. Then both branches of the if call _get_capecod_aprioris, recomputing the apriori from this X_new and the sample_weight passed in. No path directly reuses the apriori_ stored at fit time.
That the recomputation uses the fit-time CDFs, we confirmed with two run results. First, the hand calculation for row B above, plugging in the fit-time CDFs, matches chainladder's output to 10 decimal places. Second, pred.ldf_ == cc.ldf_ prints True.
Reusing the fitted ldf_ is not unusual in itself. As we noted in #1274 on September 4, Chainladder and BornhuetterFerguson do the same in predict. What sets CapeCod apart is that it also re-estimates the apriori, so the result mixes inputs from two periods.
How the maintainer sees it
#1274 was not originally about this. Its title is "CapeCod.predict infers the fitted grain from key_labels: should that inference exist?", asking whether grain inference should exist at all, and the first line of the body says "No bug here". By September 6 the discussion had reached the point where henrydingliu (one of the project's CODEOWNERS, and the person who merged our #1275) wrote:
the current implementation of capecod re-estimates
apriori_on every inference. there is essentially no model persistence. as in, i can't save the capdecod from last quarter and reapply it this quarter.
("capdecod" is the original spelling.) In plain terms: the current CapeCod re-estimates apriori_ on every inference, there is essentially no model persistence, and a CapeCod saved last quarter cannot be applied to this quarter.
Later the same day, he added a practical observation and the current state of the package:
in practice, the a priori for a BF method is often determined based on some historical chainladder result.
this package currently doesn't support putting this entire analysis into a pipeline. something we'll have to revisit in the future
In the same comment, he laid out two paths for the community to choose between:
if consensus is 'as is', then we make explicit disclaimer in the docstring around the re-estimation behavior - if consensus is 'no re-estimation', then a bigger refactor effort would be needed.
If the consensus is to keep things as they are, the docstring states the re-estimation explicitly. If the consensus is no re-estimation, a bigger refactor is needed. As for whether to warn on every call, both sides agreed that putting it in the docstring is enough. Our reason at the time: a warning that fires on every call is one people switch off.
On September 17 he opened #1385, titled "[BUG/BRK] Is CapeCod a property or an estimator?", with the core question written like this:
what is the capecod method at its core? is it just a way of calculating a bf apriori? do we expect that apriori to update? if we do expect that apriori to change, capecod ultimate actually becomes a property, rather than an estimated result.
Put another way: what is Cape Cod at its core? Is it only a way to compute a BF apriori? Do we expect the apriori to update? If we do, CapeCod's ultimate becomes a "property" rather than an estimated result.
He attached a small example to #1385. We re-ran it unchanged on 08ded050, and the output matches what the issue prints: the model fits an apriori of 0.625 on the first state, and predict on a new state with only two origin years returns 0.416667. On the same new state, Chainladder's predict gives 3750 and 2500, BF gives 4500 and 4500, and CapeCod gives 3500 and 3166.67. The example is his, not ours.
Freezing the apriori is BF, but only in-sample
Our September 6 comment also contained this line:
Persistence and re-estimation are the same fork, and taking one means opting out of the thing CapeCod does.
Persistence and re-estimation are one fork in the road: pick either side, and you have given up the thing CapeCod does. That is our framing. The maintainer did not quote it and did not say whether he agrees.
The equality behind it is not a new finding. The official user guide already says default CapeCod "can be emulated by" BornhuetterFerguson. In the source, CapeCod.fit builds BF's expectation directly and hands it to BF:
self.expectation_ = sample_weight * self.detrended_apriori_
So if you multiply detrended_apriori_ by the exposure and feed it to BF, on the same data both sides give a total ultimate of 78871.9279, with a cell-by-cell difference of 0.0. This holds by construction, and on September 13 we added ourselves: "That is close to tautological".
What matters is out of sample. Carry the stored detrended_apriori_ to the next diagonal (BF's sample_weight set to the new exposure times the stored value), and the output is:
| Item | Result |
|---|---|
Origin years covered by the stored detrended_apriori_ |
2007 to 2012 |
sample_weight for origin year 2013 |
nan |
ultimate_ for origin year 2013 |
6283.00, equal to reported losses |
ibnr_ for origin year 2013 |
NaN |
| Total ultimate, first six origin years | frozen 80242.93, CapeCod.predict 80774.45, about 0.66% apart |
The newest origin year gets no IBNR at all, and ultimate_ contains no NaN, so ultimate_ alone does not show that anything is missing. This lines up with the maintainer's correction on September 6: "detrended_apriori_ is origin-specific because the trend and olf vectors are origin-specific." The stored values are tied to the old origin years, and a new origin year has no counterpart.
The conditions on the 0.66% must be stated with it: one small public triangle, paired with a flat synthetic exposure of 20000. It illustrates the mechanism and cannot be taken as the size of the error on a real book.
Three questions to ask before calling predict
This example applies to any model that is fitted once and then used for repeated predict calls, actuarial or not:
- Which quantities are stored, and which are recomputed on every call? CapeCod stores
ldf_and recomputesapriori_. If the documentation does not make this clear, print the key attributes from fit and from predict side by side and compare. - Which quantity accounts for the difference between predict and a clean refit? Run each once on the same new data and compare item by item. In this example the whole difference comes from the CDF; the losses and the exposure are identical.
- Is the thing you want to persist defined for the new data? A parameter indexed by origin year has no value when a new origin year arrives. In this example the consequence is zero IBNR for the newest origin year, with no error raised. A development pattern indexed by line of business has the same gap: when the prediction data carries a line of business the model never saw, Chainladder's predict drops those rows without raising.
Until the community reaches a conclusion on #1385, if your workflow is "apply last quarter's CapeCod to this quarter", at minimum know that the apriori_ predict returns is no longer last quarter's.
Current status (2026-09-26)
- #1274: open, 14 comments. The last one is the maintainer's, from September 17: "new issue raised at #1385. thanks for all the help. we'll be able to finally close this out soon after some community input."
- #1385: open, label Triage Pending, zero replies.
- #1306: a separate PR of ours. It deals with a warning during grain inference, not with re-estimation. It is not merged, is awaiting review, and has conflicts with main.
No direction has been decided.
Sources
- casact/chainladder-python issue #1274 (opened by us). Comments quoted: maintainer 2026-09-06 17:58 UTC (comment 5561081469), ours 2026-09-06 18:58 UTC (comment 5561426764), maintainer 2026-09-06 20:55 UTC (comment 5562110562), ours 2026-09-13 (comment 5652358664), maintainer 2026-09-17 (comment 5720823903). As checked on 2026-09-26, none of the 14 comments had been edited since posting.
- Issue #1385 (opened by henrydingliu, 2026-09-17).
- Official user guide, Methods; corresponding source file
docs/user_guide/methods.ipynb, cells 28 and 29. - Source code:
CapeCod.fitandCapeCod.predictinchainladder/methods/capecod.py, commit08ded050188728f155f6ad71b7d7b7586da6eea2(main, 2026-09-25). Theifinpredictis line 329 at this commit and line 324 in PyPI 0.10.1; thecapecod.py:325cited in the #1274 body is from an earlier commit. - Reproduction: dataset
cl.load_sample("ukmotor"),trend=0, flat exposure of 20000, code as in the snippet above. Environment: a standalone venv on Python 3.12.4 (py -3.12 -m venv venv), numpy 2.5.3, pandas 3.0.6. A plainpip install chainladder==0.10.1is enough to reproduce. Main and three other commits were run withPYTHONPATHpointing at that commit's source, using the same code:91a942f1(the #1275 merge point),cd24e1d2(main at the time of the September 6 comment) and6b51f649(the commit mentioned in the September 13 comment). Every number is the same across all five versions. - Hand rebuild of the mechanism and the #1385 example: same environment, commit
08ded050, run on 2026-09-26.
FAQ
Does CapeCod.predict reuse the apriori computed at fit time?
No. Every predict call recomputes the apriori from the new triangle and new exposure passed in, and reuses only the fit-time development factors, ldf_. In the ukmotor example, the fitted apriori is 0.6572660657 and predict returns 0.6894525318.
Then is the predict result the same as refitting on the new data?
Not that either. A clean refit on the new data gives 0.7062253126. predict and the refit use exactly the same losses (latest diagonal total 75672) and exposure. The only difference is the development pattern: predict uses the prior period's cumulative development factors, the refit uses new ones. For example, for origin year 2008 at 72 months, predict uses 1.000000 and the refit uses 1.027530.
Is this a bug in chainladder?
When we opened #1274, the first line of the body said No bug here, raising it as a design question. On 2026-09-17 maintainer henrydingliu opened a separate issue, #1385 (titled Is CapeCod a property or an estimator?), asking the community to decide whether CapeCod should stay as it is with the docstring stating that it re-estimates, or change to not re-estimating (which would need a bigger refactor). As of 2026-09-26, #1385 has no replies.
Is freezing CapeCod's apriori and reusing it the same as BornhuetterFerguson?
Only in-sample. The official user guide already says default CapeCod can be emulated by BornhuetterFerguson, and on the same data both give a total ultimate of 78871.9279. But carry the stored detrended_apriori_ to the next diagonal and the new origin year 2013 has no corresponding value, so its ultimate equals the reported losses of 6283.00, with no IBNR booked at all.
How big is this gap in practice?
We have no data to answer that. The only thing we measured is ukmotor, a small public triangle, with a flat 20000 exposure copied from the docstring example: the total ultimate for the first six origin years differs by about 0.66%. It illustrates the mechanism and does not represent the reserve error on a real book.