FastRiskScore: Fitting sparse integer risk scores via Autoresearch
Chandan Singh · October 2026
A risk score is a short list of conditions, each worth a few integer points. You add the points of the conditions that apply and look up the risk for the total in a table. Clinical scores often work this way, so a doctor can compute and check them by hand. Here is an example that predicts whether a breast tumour is benign:
- worst area ≤ 906.6+3
- worst concave points ≤ 0.1479+3
- mean texture ≤ 20.99+2
Choosing the conditions and points is a slow integer optimization problem. We used autoresearch to build FastRiskScore, which builds on prior work in the field like RiskSLIM and FasterRisk. On held-out datasets FastRiskScore generally fits much faster than competing baselines, without sacrificing training loss or test AUC.
1. Quickstart & example usage
from imodels import FastRiskScoreClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
model = FastRiskScoreClassifier(k=5).fit(X_train, y_train)
proba = model.predict_proba(X_test)
k(default 5): the most conditions the score may use.max_points(default 5): the largest number of points one condition can carry.n_thresholds(default 9): thresholds per numeric column, so the default splits each at its deciles. With more than 19 (for example 99, the percentiles) the search switches to settings tuned for fine thresholds.binarize(default True): numeric columns becomex ≤ tconditions and categorical columns one condition per level, so each condition refers to one of the original columns. Set it to False to use binary or small-integer columns as they are.time_limit(default 60 s): the search returns the best score found by then.
The fitted model: points, total score and risk curve
Here is the fitted model, fit in 0.02 seconds:
- worst area ≤ 906.6+5
- worst texture ≤ 20.14+4
- worst concave points ≤ 0.1479+4
- mean texture ≤ 20.99+3
- worst concave points ≤ 0.09804+2
The fitted model stores the score and its risk curve:
model.points_ # {"worst area <= 906.6": 5, ...}, the lines of the score
model.total_score(X) # the total points of each row
model.risk(score) # 1 / (1 + exp(-(model.scale_ * score + model.intercept_)))
model.train_loss_ # the mean training log loss of that risk
The points determine the total score of each row. The risk for a total comes from a logistic
curve, risk(s) = 1 / (1 + exp(-(a·s + b))), with a =
model.scale_ and b = model.intercept_; both are fit to the
training data once the points are chosen. The points are chosen to make the log loss of this
curve as small as possible (Sec 3), so the risks in the diagram are
calibrated on the training data. RiskSLIM and FasterRisk show the same kind of table; they write
b as an integer offset and 1/a as a multiplier.
2. Experimental results
The autoresearch loop. FastRiskScore came out of the autoresearch loop of
Agentic-imodels (code on
github).
Two coding agents (Claude Code with Opus 5.5) started from the same solver, FasterRisk rewritten as a single Python file, and used the
same benchmark. One was asked to lower the loss and the other to
make the solver faster. In each round an agent changed the solver, ran the benchmark, and kept
the change only if the loss went down without the time rising by more than 10%, or the time
went down by at least 10% without the loss rising. The agents ran for about 75 hours of autoresearch. The
shipped version comes from later loops (below); Fig 1 and
Table 1 show it.
All solvers are scored the same way. A solver returns integer points in [−5, 5] for at most
k of a dataset's binary features. The benchmark then fits the risk curve for those
points itself and records 3 numbers: the training log loss of that curve, which is the
quantity being minimized; the test AUC of the total score on a held-out 20% split; and the fit
time.
Datasets, baselines, and evaluations. The benchmark that the agents optimize over (which we call the "visible set"),
has 70 problems: 14 datasets × 5 values of k (3, 4, 5, 7 and 10). Each problem
has a 60 s limit and runs on one core.
We compare against 19 baselines that run without a commercial solver, from FasterRisk and exact
solvers to sparse models with rounding and recent scoring-system packages (listed below).
The 14 visible datasets (from the RiskSLIM and FasterRisk papers, and OpenML)
| dataset | what it is | rows | binary features | positive rate |
|---|---|---|---|---|
adult | UCI Adult income, RiskSLIM binarization | 12,500 | 36 | 0.24 |
bank | UCI Bank telemarketing, RiskSLIM binarization | 12,500 | 57 | 0.11 |
breastcancer | UCI Wisconsin breast cancer, 9 cytology scores (1 to 10) | 683 | 9 | 0.35 |
mammo | UCI Mammographic mass, RiskSLIM binarization | 961 | 14 | 0.46 |
mushroom | UCI Mushroom, one-hot | 8,124 | 113 | 0.48 |
spambase | UCI Spambase, 57 word and character frequencies (real-valued) | 4,601 | 57 | 0.39 |
fico | FICO HELOC credit (OpenML 45023) | 10,000 | 142 | 0.50 |
compas | ProPublica COMPAS two-year recidivism (OpenML 42192) | 5,278 | 32 | 0.47 |
australian | Statlog Australian credit | 690 | 79 | 0.45 |
heart | Statlog heart disease | 270 | 57 | 0.44 |
ionosphere | UCI Ionosphere radar returns | 351 | 267 | 0.36 |
ilpd | Indian liver patient records | 583 | 79 | 0.29 |
magic | MAGIC gamma telescope | 12,500 | 90 | 0.35 |
haberman | Haberman breast cancer survival | 306 | 27 | 0.27 |
The first six are the files distributed with RiskSLIM,
which FasterRisk also uses; breastcancer and spambase keep real-valued columns, and mammo has
one small-integer column. The other eight are binarized here: each numeric column becomes
indicators x ≤ t at its deciles, each categorical column one indicator per
level. Datasets above 12,500 rows are subsampled to 12,500, and every dataset is split 80/20
into training and test rows.
The 27 hidden datasets (from TabArena)
| dataset | rows / full size | columns | binary features | positive rate |
|---|---|---|---|---|
Amazon_employee_access | 12,500 / 32,769 | 9 | 107 | 0.06 |
APSFailure | 12,500 / 76,000 | 170 | 399 | 0.02 |
Bank_Customer_Churn | 10,000 / 10,000 | 10 | 54 | 0.20 |
Bioresponse | 3,751 / 3,751 | 1,776 | 306 | 0.46 |
blood-transfusion | 748 / 748 | 4 | 32 | 0.24 |
churn | 5,000 / 5,000 | 19 | 148 | 0.14 |
coil2000_insurance_policies | 9,822 / 9,822 | 85 | 176 | 0.06 |
credit-g | 1,000 / 1,000 | 20 | 87 | 0.30 |
credit_card_clients_default | 12,500 / 30,000 | 23 | 154 | 0.22 |
customer_satisfaction_in_airline | 12,500 / 129,880 | 21 | 109 | 0.45 |
Diabetes130US | 12,500 / 71,518 | 47 | 196 | 0.09 |
E-CommereShippingData | 10,999 / 10,999 | 10 | 52 | 0.40 |
Fitness_Club | 1,500 / 1,500 | 6 | 41 | 0.30 |
GiveMeSomeCredit | 12,500 / 150,000 | 10 | 57 | 0.07 |
hazelnut-contaminant | 2,400 / 2,400 | 30 | 270 | 0.50 |
HR_Analytics | 12,500 / 19,158 | 12 | 98 | 0.25 |
in_vehicle_coupon | 12,500 / 12,684 | 24 | 112 | 0.43 |
Is-this-a-good-customer | 1,723 / 1,723 | 13 | 75 | 0.11 |
kddcup09_appetency | 12,500 / 50,000 | 212 | 294 | 0.02 |
Marketing_Campaign | 2,240 / 2,240 | 25 | 140 | 0.15 |
NATICUSdroid | 7,491 / 7,491 | 86 | 80 | 0.35 |
online_shoppers_intention | 12,330 / 12,330 | 17 | 99 | 0.15 |
polish_companies_bankruptcy | 5,910 / 5,910 | 64 | 387 | 0.07 |
qsar-biodeg | 1,054 / 1,054 | 41 | 208 | 0.34 |
seismic-bumps | 2,584 / 2,584 | 15 | 66 | 0.07 |
taiwanese_bankruptcy | 6,819 / 6,819 | 94 | 360 | 0.03 |
jm1 | 10,885 / 10,885 | 21 | 167 | 0.19 |
Every binary classification dataset of TabArena (OpenML suite
tabarena-v0.1) except bank-marketing, heloc and diabetes, which the visible set
already contains. They were binarized by the same rule as the visible set, with at most 40
columns kept (those with the highest mutual information with the label on the training split),
and never used during development. Builder: evolve_slim/data/build_data.py.
The 19 baselines
| baseline | family | source and how it is run |
|---|---|---|
| FasterRisk | FasterRisk | Liu et al. 2022; package fasterrisk 0.1.10, default settings |
| FasterRisk, wide search | FasterRisk | the same with a beam of 50 × 50, a pool of 200, 100 swaps and 100 multipliers |
| RiskSLIM | exact | Ustun & Rudin 2019; risk-slim with the free CPLEX Community Edition |
| Cutting planes, HiGHS | exact | RiskSLIM's cutting-plane algorithm reimplemented with the open-source HiGHS solver |
| SLIM | exact | Ustun & Rudin 2016; its 0-1 loss integer program solved with HiGHS |
| abess + rounding | sparse + rounding | Zhu et al. 2022; best-subset selection, package abess |
| fastSparse + rounding | sparse + rounding | Liu et al. 2022; the L0 path, package fastsparsegams |
| OKRidge + rounding | sparse + rounding | Liu et al. 2023; optimal sparse ridge support, vendored |
| L1 path + rounding | sparse + rounding | supports along the L1 logistic path (scikit-learn) |
| L0Learn + rounding | sparse + rounding | Dedieu et al. 2021; L0L2 logistic path with swaps, R package L0Learn |
| OKGLM + rounding | sparse + rounding | Liu et al. 2025; certified k-sparse logistic regression in [−5, 5], vendored, CPU |
| skscope + rounding | sparse + rounding | Wang et al. 2024; k-sparse logistic regression by splicing, package skscope |
| Probabilistic scoring list | scoring system | Hanselle et al. 2025; package scikit-psl, vendored |
| AutoScore | scoring system | Xie et al. 2020; reimplemented (R only) |
| riskscores, annealing | scoring system | R package riskscores 1.3.0 |
| riskscores, coordinate descent | scoring system | the same package, its coordinate-descent method |
| Rounded L1 logistic | common practice | L1 logistic regression with k features, scaled to a largest point of 5 and rounded |
| Unit weighting | common practice | the L1-selected features, each worth +1 or −1 by its sign |
| imodels SLIMClassifier (before) | common practice | imodels 3.0.2 without a commercial solver: a rounded L2 logistic regression |
The baselines came from two searches for methods that can be run without a
commercial solver: one of packages for risk scores and sparse logistic regression, and a later one of
2023 to 2026 papers that cite FasterRisk, RiskSLIM, fastSparse, OKRidge, scoring lists or the
certifier of Liu et al. (ICML 2025). Each "+ rounding" method gives a sparse real-valued solution with
at most k features, which FasterRisk's rounding step turns into points. Penalized
methods (riskscores, L0Learn, the L1 path, imodels' SLIMClassifier) have their penalty searched until
at most k points are nonzero, and points are clipped to [−5, 5]. scikit-psl and
okridge pin numpy < 2 and are vendored with small fixes (a numpy 2 incompatibility in
scikit-psl, a call to a removed method in okridge), and scikit-psl gains a stop after k
stages. AutoScore is only available in R, so its recipe (rank by random-forest importance, fit a
logistic regression, divide by the smallest coefficient and round) is reimplemented on the binary
features. RiskSLIM has no open-solver version, so its cutting-plane algorithm is reimplemented with
HiGHS, re-solving the integer program each round and starting from rounded heuristic solutions as
RiskSLIM does; it proves optimality on 6 of the 70 visible problems. L0Learn runs from its R package,
since the Python package has no build past Python 3.10, and the R methods run in one persistent R
process per worker so that R's start-up is not timed. OKGLM has no intercept, so a constant column is
added. imodels' previous SLIMClassifier solves its integer program only with a commercial solver
installed; without one, which is the default, it rounds an L2 logistic regression. Probabilistic
scoring lists are killed at 3× the time limit on 10 visible and 51 hidden problems, and
riskscores often passes the time limit, since it checks the clock only between fits; a solver
without a score is charged the base rate. Left out: scorepyo (abandoned, pinned to 2022 packages),
weight-of-evidence scorecards (optbinning, scorecardpy; not sparse integer points), GroupFasterRisk
(the same as FasterRisk without feature groups), recent methods without public code (a GPU branch
and bound for k-sparse GLMs, net-benefit risk scores) and those that optimise a different objective
(boosted risk scores, SHAPoint, AgentScore).
FastRiskScore is fast and accurate. Fig 1 plots time per problem against the training loss (Fig 1A-C) and the test error (Fig 1D). In all cases, FastRiskScore achieves the best performance (lower left). The loss is shown as the excess over the lowest loss any solver reached on each problem, averaged over problems. A solver that returns no score on a problem is charged the loss of predicting the base rate there.
- Autoresearch run 1 (lower loss)
- Autoresearch run 2 (faster)
- Discarded attempts
- Baselines
- FastRiskScore
Visible set (14 datasets × 5 values of k, 60 s per problem)
Hidden set (27 datasets × 5 values of k, 60 s per problem)
Hidden set, full size (k = 5, 10 min per problem)
Hidden set, test error (27 datasets × 5 values of k, 60 s per problem)
Fig 1. Training loss and test error against fit time. FastRiskScore (star)
is about 180× faster than FasterRisk on the visible set and 225× on the held-out sets (by
the median), and has a lower loss on 55 of 70, 94 of 135 and 21 of 27 problems (a higher one on 0, 8
and 0), and has the lowest test error of the integer scores. The closest baselines are sparse
models rounded with FasterRisk's rounding step. A: the visible set
that the loop was scored on, with every version either run tried (discarded attempts are faded).
B: the same protocol on 27 held-out TabArena datasets, for the versions each run kept.
C: the held-out datasets at full size, up to 120,000 training rows, at k = 5
with 10 minutes per problem. D: the test error (1 − AUC) of the solvers in B, averaged over the 135
held-out problems; the axis stops at 0.275, so the solvers above it (RiskSLIM 0.432, SLIM
0.358, cutting planes 0.333 and scoring lists 0.319) are not shown. Each point is one solver; the x-axis is its geometric-mean fit time, and in A
to C the y-axis is its mean excess training loss over the best loss any solver reached on each
problem (log scale). Large points are the mean of three runs. In C no other solver shown beats FastRiskScore
on any dataset, so its star sits at the bottom of the axis. All runs
used one core per problem on a Linux server (two Xeon Platinum 8468V), with 14 problems at a
time (27 in C).
Extra details on experimental settings
Scoring. The benchmark rejects any answer whose points are not integers, fall outside
[−5, 5] or use more than k features; no solver in the figure had such an answer.
Given the points, the risk curve is fit by Newton's method on the two numbers a and
b, a convex problem. Each solver process first fits a tiny warm-up problem, untimed, so
numba compilation is not counted; FastRiskScore's first fit on a new machine takes about 2 minutes
longer for that reason.
Best known losses. The excess loss is measured against the lowest loss any integer solver reached on the problem in any of these runs, including FasterRisk with every search width raised (a beam of 50 × 50, a pool of 200 and 100 multipliers) and a 20-minute limit, which takes about 12× its default time.
RiskSLIM. The free CPLEX Community Edition refuses problems above 1,000 variables or constraints, and RiskSLIM's cutting planes pass that limit on wider datasets; it returned no score on 25 of 70 visible problems and 102 of 135 hidden ones, and is charged the base rate there. With a full CPLEX licence it would do better, but within these time limits it still certifies optimality only on small problems, as the FasterRisk paper also found. RiskSLIM and SLIM run until their time limit on most problems, so their times sit near 60 s.
C. One run per solver at k = 5 on each of the 27 datasets at full size, 27 at a
time, with a 10-minute limit (killed at 30 minutes). RiskSLIM stops at the CPLEX Community
Edition's size limit on 22 datasets; SLIM passes its limit on 25, since HiGHS finishes its
current step before stopping, and on the remaining 2 returns a proven optimum of its 0-1
loss.
Table 1: loss, test AUC and time of every solver on the three suites
Table 1. Mean training log loss, mean test AUC and geometric-mean fit time (s) of every solver on the three suites of Fig 1. Eight of the baselines (cutting planes, abess, fastSparse, OKRidge, the L1 path, scoring lists, AutoScore and unit weighting) were added after a search for baselines that can be run, and five more (L0Learn, OKGLM, skscope and two versions of riskscores) after a search of papers that cite them (details). The last row is not a risk score: it is FasterRisk's real-valued solution before rounding, shown as a reference.
| Visible (70 problems) | Hidden (135 problems) | Hidden, full size (27) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| solver | loss | AUC | time | loss | AUC | time | loss | AUC | time |
| FastRiskScore (ours) | 0.3558 | 0.842 | 0.011 | 0.3269 | 0.794 | 0.015 | 0.3279 | 0.792 | 0.020 |
| FasterRisk | 0.3603 | 0.841 | 1.51 | 0.3272 | 0.793 | 3.28 | 0.3282 | 0.790 | 4.82 |
| FasterRisk, wide search | 0.3582 | 0.841 | 18.2 | 0.3271 | 0.793 | 47.3 | – | – | – |
| RiskSLIM (CPLEX CE) | 0.4524 | 0.718 | 13.6 | 0.4046 | 0.568 | 9.98 | 0.4031 | 0.560 | 16.2 |
| Cutting planes, HiGHS (RiskSLIM's method) | 0.4097 | 0.776 | 57 | 0.3739 | 0.667 | 60 | 0.3602 | 0.695 | 600 |
| SLIM (HiGHS) | 0.4371 | 0.767 | 49.3 | 0.3895 | 0.642 | 60.2 | 0.3835 | 0.647 | 542 |
| abess + sequential rounding | 0.3731 | 0.842 | 0.207 | 0.3303 | 0.789 | 0.512 | 0.3317 | 0.788 | 0.790 |
| fastSparse + sequential rounding | 0.3738 | 0.836 | 0.123 | 0.3303 | 0.791 | 0.207 | 0.3304 | 0.790 | 0.304 |
| OKRidge + sequential rounding | 0.3654 | 0.841 | 20.4 | 0.3286 | 0.790 | 26.5 | 0.3297 | 0.786 | 214 |
| L1 path + sequential rounding | 0.4087 | 0.813 | 0.099 | 0.3546 | 0.744 | 0.231 | 0.3581 | 0.744 | 0.370 |
| Probabilistic scoring list | 0.4185 | 0.787 | 13.8 | 0.3623 | 0.681 | 51.2 | 0.3366 | 0.763 | 127 |
| L0Learn + sequential rounding | 0.3633 | 0.843 | 0.880 | 0.3290 | 0.792 | 2.84 | 0.3293 | 0.790 | 4.21 |
| OKGLM (certified k-sparse) + sequential rounding | 0.3609 | 0.842 | 25.3 | 0.3293 | 0.792 | 31.4 | 0.3296 | 0.793 | 175 |
| skscope + sequential rounding | 0.3751 | 0.837 | 0.222 | 0.3318 | 0.791 | 0.720 | 0.3325 | 0.792 | 1.12 |
| riskscores (coordinate descent) | 0.4476 | 0.757 | 12.7 | 0.3817 | 0.666 | 36.6 | 0.3713 | 0.699 | 111 |
| riskscores (annealing) | 0.4876 | 0.701 | 38.3 | 0.4097 | 0.585 | 70.1 | 0.3821 | 0.665 | 283 |
| AutoScore | 0.4000 | 0.821 | 0.260 | 0.3461 | 0.763 | 0.495 | 0.3469 | 0.768 | 0.869 |
| Rounded L1 logistic | 0.4096 | 0.812 | 0.068 | 0.3547 | 0.743 | 0.180 | 0.3580 | 0.744 | 0.292 |
| Unit weighting | 0.4311 | 0.804 | 0.062 | 0.3665 | 0.732 | 0.180 | 0.3665 | 0.739 | 0.293 |
| imodels SLIMClassifier (before) | 0.4260 | 0.789 | 0.258 | 0.3583 | 0.736 | 0.810 | 0.3612 | 0.732 | 1.42 |
| Real-valued k-sparse (not integer) | 0.3574 | 0.844 | 0.867 | 0.3265 | 0.795 | 1.56 | 0.3275 | 0.794 | 2.27 |
Table 2: six later loops on versions of both suites with 99 thresholds per numeric column instead of 9
The suites above split each numeric column at its 9 deciles. We also built versions of the visible
and hidden suites that split each numeric column at its 99 percentiles, which gives the search many
more features to choose from (990 instead of 90 on magic). The six RiskSLIM datasets were rebuilt
from their raw OpenML sources for this. Six more loops ran on the fine visible suite with the same keep rule. Loops 1 and 2 both
started from FastRiskScore; each later loop started from the best version of the one before. From loop 3 on, a change was also
rejected if it lowered the mean test AUC by more than 0.001. The fine hidden suite (27 datasets
× 5 values of k) was not used during the loops. Table 2
gives the best version of each loop that kept a change, on both.
Table 2. Mean training log loss, mean test AUC and geometric-mean
fit time (s) on the fine suites (99 thresholds per numeric column). Run 2's version is the first
release of FastRiskScore; each loop row is that loop's best version and its main change; the last
row is the shipped solver with n_thresholds > 19.
| Fine visible (70) | Fine hidden (135) | |||||
|---|---|---|---|---|---|---|
| solver | loss | AUC | time | loss | AUC | time |
| FasterRisk | 0.34282 | 0.848 | 2.16 | 0.32448 | 0.792 | 4.62 |
| Run 2's version (first release) | 0.34119 | 0.850 | 0.068 | 0.32416 | 0.792 | 0.153 |
| Loop 1: wider local search with kicks | 0.33882 | 0.847 | 0.259 | 0.32393 | 0.793 | 0.322 |
| Loop 2: thresholds of one column handled together | 0.33982 | 0.848 | 0.032 | 0.32400 | 0.793 | 0.050 |
| Loop 3: minimum support per feature | 0.33998 | 0.851 | 0.023 | 0.32416 | 0.793 | 0.042 |
| Loop 4: faster | 0.33996 | 0.851 | 0.011 | 0.32414 | 0.793 | 0.021 |
| Loop 5: exact polish of the best score | 0.33976 | 0.851 | 0.011 | 0.32414 | 0.793 | 0.021 |
| FastRiskScore (shipped, fine settings) | 0.34079 | 0.852 | 0.011 | 0.32414 | 0.794 | 0.020 |
Versions: loop 1 f31_fastsame, loop 2 g24_fastcol, loop 3
h8_fastbeam, loop 4 i30_final, loop 5 j18_poliso. Hidden-suite AUCs
to four decimals: FasterRisk 0.7917, run 2's version 0.7924, loops 1 and 2 0.7927, loop 3 0.7930, loops
4 and 5 0.7932, shipped 0.7937. Times vary by about 5% between runs; the loop 1 and 2 hidden rows were run while
another loop used the same machine.
- Loss. Every loop beats FastRiskScore on the fine visible suite, but on the fine hidden suite the gains are small: 0.0002 at most. Loop 1 has the lowest loss on both suites, at 6 to 15 times the hidden-suite time of the later loops.
- Feature handling. Loop 2 found that the 99 indicators of a numeric column are nested
(
x ≤ tfor increasingt), and handled each column's indicators as one ordered variable. This made the search faster and let the beam pick the best threshold of a column instead of filling with neighbouring thresholds of one column. - AUC. Loop 3 required each chosen indicator to cover at least √N training rows on each side. Without it, the search picks thresholds that fit a handful of training rows. Against loop 2, it raised visible AUC by 0.003 for a loss increase of 0.0002, and the AUC gain held on the hidden suite (0.7927 to 0.7930).
- Speed. Loop 4 reached about the same loss as loop 3 in half the time: 0.021 s on the hidden suite, 7× faster than FastRiskScore and 220× faster than FasterRisk.
- Saturation. Loop 5's only kept change, an exact polish of the final score, lowered the
visible loss by 0.0002. Almost all of that came from two near-separable problems (ionosphere at
k= 7, mushroom atk= 10), and it changed 2 of the 135 hidden problems. On the problems where the polish runs, no single change of one feature's points, and no swap of one feature, lowers the loss of the final score. A sixth loop (30 attempts: new starting points, tabu walks, pair moves) kept nothing; running its whole search 4 times longer lowered the mean loss by 0.00001.
Table 3: how the shipped version was found (two loops on the decile suites), compared with run 2's version
The fine-track versions are 3× faster than run 2's version, but on the decile suites they
lost 0.0014 in visible loss, mostly on spambase and ionosphere. Two causes accounted for most of
it: loop 3's minimum support per feature, and beam rules that kept one threshold per variable and
one child per set of variables, which removed good feature sets on decile data. A loop on the
decile visible suite started from loop 5's version with the minimum support switched off and kept
four changes: candidates whose real-valued coefficient rounds to 0 are ranked last in the beam
(spambase, whose large-scale columns collapsed 3 of 5 final feature sets to 2 features), the
search kernels are compiled before the first fit (about 4 s had been spent compiling inside
a few fits), the beam diversity rules are switched off, and the beam is trimmed back to 8 parents.
A second loop from that version found nothing that held on the hidden suite. The package ships the second
loop's final version (n17_lean, with the same hidden-suite results), and switches to the
fine-track settings when n_thresholds > 19. Table 3 compares it with
run 2's version.
Table 3. Mean training log loss, mean test AUC and geometric-mean fit time (s) of the shipped version and run 2's version on the suites of Table 1.
| Visible (70 problems) | Hidden (135 problems) | Hidden, full size (27) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| solver | loss | AUC | time | loss | AUC | time | loss | AUC | time |
| FasterRisk | 0.36029 | 0.841 | 1.49 | 0.32720 | 0.793 | 3.27 | 0.32823 | 0.790 | 4.82 |
| Run 2's version (first release) | 0.35592 | 0.842 | 0.035 | 0.32695 | 0.794 | 0.057 | 0.32789 | 0.793 | 0.079 |
| FastRiskScore (shipped) | 0.35583 | 0.842 | 0.011 | 0.32688 | 0.794 | 0.015 | 0.32789 | 0.792 | 0.020 |
AUCs to four decimals, visible / hidden / full size: run 2's version 0.8424 / 0.7937 / 0.7926, shipped 0.8421 / 0.7939 / 0.7920. Times vary by about 10% between runs.
On the hidden suite the shipped version has a lower loss than run 2's version on 23 of the 135 problems and a higher one on 10, and it is 3 to 4× faster on the three suites. Test AUC is within 0.0006 of run 2's version on every suite.
3. Methods: FastRiskScore
The problem is the one RiskSLIM and FasterRisk solve. Over integer points w with at most k nonzero entries, each in [−5, 5], minimise
with \(\tilde y_i \in \{-1, +1\}\). The inner minimum over a and b is the risk curve of Sec 1. FasterRisk first finds good real-valued coefficients and then rounds them, trying a few multipliers 1/a with the intercept fixed by the rounding. FastRiskScore keeps FasterRisk's first stage, rewritten to run faster, and replaces the rounding stage with a search over the integer points, where each candidate is scored by the loss above with a and b fit again. The shipped version has five steps:
- Compress the data. Identical rows (same features and label) are merged into one row with
a count, and the nested thresholds of one numeric column (
x ≤ tfor increasingt) become one ordered variable, so the counts for every threshold come from one table. - Beam search for real-valued solutions (FasterRisk's): grow the set of features one at a time from the 8 best sets. Each set proposes its 10 most promising features, ranked by a one-step Newton bound, at most one per threshold band of a variable, and the 20 best proposals are fit by Newton's method on rows grouped by their values on the chosen features. A set whose coefficient on a raw numeric column would round to 0 is ranked last.
- Round each of the 8 best final solutions at 20 scales, and at 4 more where the largest points are clipped at the bound so the others keep finer ratios, keeping the rounding with the lowest loss after refitting a and b.
- Local search on the integer points from the best roundings: change one feature's points or add a feature, move a threshold along its column, then swap one feature for another; finally try the best swaps followed by a real-valued refit and rounding. The loss of a move is computed from a table of (score, feature value) counts; an estimate ranks the moves and the best are checked exactly.
- Exact polish when the score nearly separates the classes: every single change is checked exactly, skipping candidates whose lower bound cannot beat the current loss.
With more than 19 thresholds per column the search also requires each chosen feature to cover at least √n rows on each side, keeps one threshold per variable in the beam, and uses 10 sets in the beam and 5 roundings; these settings come from the fine-threshold loops of Sec 2.
Loop findings
Searching the integer points directly, scored by the refit loss (both runs)
Both runs found this change first, and it lowered the loss more than any other: on the visible set
it moved the mean excess loss from about +0.002 to −0.001 (measured against the original best
known losses) in each run. FasterRisk judges a rounding by its loss at a fixed multiplier, but a
risk score is used with its fitted curve, so a rounding that is worse at the fixed multiplier can
be better once a and b are fit again. On spambase at k = 3, FasterRisk's
rounding sets one of its three features to zero, because at that feature's scale the best integer
point rounds to 0. The local search puts the feature back and lowers the loss by 0.095.
At first the local search added 60 to 80% to the run time, because every candidate move required
fitting a and b again over all rows. Two changes made it cheap. The loss is computed
from the counts of positive and negative rows at each value of the total score (with 5 features
the total takes at most 51 values), and moves are ranked by an estimate so that only the best two
are checked exactly. Run 2 found that its first estimate was too optimistic for moves that change
the spread of the scores (0.58 estimated against 0.63 true on compas at k = 3). It
fixed this by recomputing the cells a move touches after one Newton step in (a,
b).
Speed changes
A profile of FasterRisk on bank at k = 10 (19.9 s) showed 82% of the time in
coordinate descent, which reads every row for each coordinate step. Merging identical rows and
fitting by Newton's method on the few chosen features made both runs 5× faster. Run 2 then
grouped each beam candidate's rows by their values on its features (another 2.5×), moved the
creation of each new candidate into one compiled function together with smaller changes
(−34% time), and computed the gradient and swap screens as single matrix products
(−14%, then −12%). After these changes each step reads the data once instead of once
per candidate.
The integer search also made parts of FasterRisk's first stage unnecessary. Run 2 removed FasterRisk's diverse pool (swapping each feature for up to 50 others) and started the local search from the 10 beam solutions instead, which gave a lower loss. Run 1 kept the pool but reduced it from 50 to 10 solutions and the rounding multipliers from 20 to 5, with no change in loss.
Changes that didn't yield improvements
- Local search from an empty score, without the beam search: much worse (excess loss +0.010 against −0.001), so the features chosen by the beam search matter.
- Cheaper ways to rank beam candidates before fitting them (a one-step score test, a one-coordinate fit): all raised the loss by 0.0002 to 0.0007.
- Smaller beams: 5 parents instead of 10 raised the loss by 0.003.
- Random restarts and perturbations of the local search: large gains on a few problems where it
got stuck (mushroom at
k= 7), but less than 0.0001 on average, for 2 to 3 times the time. 300 random restarts per problem on 15 problems beat the final version on only 4. - More starts, deeper swaps, rescaling all points at once, refitting the best set of features and rounding again at several resolutions: gains of at most 0.0002, for 20 to 60% more time, so none passed the keep rule.
Full table of kept versions
| version | what changed | mean loss | time (s) | test AUC |
|---|---|---|---|---|
| Run 1 (lower loss) | ||||
pyfasterrisk_v1 | FasterRisk in one file: beam search, diverse pool, star-ray search with sequential rounding | 0.36029 | 1.444 | 0.8406 |
v3_fastnewton | the same algorithm on unique (x, y) rows with counts, numba kernels, Newton on the support in place of coordinate descent | 0.35920 | 0.279 | 0.8415 |
v5_tunetop | beam: fine-tune only the children with the best one-coordinate fit | 0.35924 | 0.211 | 0.8424 |
v6_numbaray | star-ray search and sequential rounding in one numba kernel | 0.35924 | 0.176 | 0.8424 |
v9_pool10 | integer local search scored by the calibrated loss, on score histograms; rounding pool 50 → 10 | 0.35661 | 0.163 | 0.8427 |
v11_pooltune | diverse pool: fine-tune only the most promising swaps | 0.35661 | 0.143 | 0.8427 |
v24_child5 | rescale and paired ±1 moves to polish the best result; beam child size 10 → 5 | 0.35631 | 0.142 | 0.8426 |
v27_cheap1d | ranking fits capped at 3 Newton steps; 20 swap attempts per feature | 0.35632 | 0.122 | 0.8425 |
v28_ray5 | 5 multipliers in the star-ray search instead of 20 | 0.35632 | 0.105 | 0.8425 |
v29_tune1 | row deduplication by a random projection; fewer fine-tuned candidates | 0.35636 | 0.087 | 0.8427 |
v34_2starts | local search in numba on the support; 2 starts instead of 3 | 0.35640 | 0.074 | 0.8431 |
v36_robust | end of run: same points, plus a deadline and kernels compiled at import | 0.35640 | 0.071 | 0.8431 |
| Run 2 (faster) | ||||
v2_numba | FasterRisk in numba on unique (x, y) rows; projected Newton in place of coordinate descent | 0.35981 | 0.296 | 0.8414 |
v5_nopool | integer local search scored by the calibrated loss (value changes, additions, swaps); diverse pool dropped | 0.35637 | 0.264 | 0.8427 |
v7_groups | beam nodes keep rows grouped by (y, support pattern): at most a few hundred groups instead of all rows | 0.35637 | 0.107 | 0.8427 |
v14_numbachild | faster data setup, swap screen, score binning by counting, visited-state memory; child creation in one numba call | 0.35639 | 0.071 | 0.8428 |
v16_swaplast | value changes and additions before swaps; swaps only when those fail | 0.35622 | 0.063 | 0.8427 |
v17_calround | rounding by the calibrated loss over 20 scales, replacing star-ray sequential rounding | 0.35602 | 0.066 | 0.8430 |
v18_roundnb | that rounding in one numba kernel | 0.35597 | 0.057 | 0.8430 |
v24_blasscreen | better move estimate (touched cells exact after a Newton step in (a, b)); swap screen as two BLAS products | 0.35592 | 0.049 | 0.8428 |
v30_lazycodes | column codes computed lazily; exact checks skipped when the estimate rules a move out | 0.35592 | 0.042 | 0.8428 |
v34_gemm | swap screens for all k removals, and beam gradients for all parents, as one matrix product each | 0.35592 | 0.037 | 0.8424 |
v35_scratch | run 2's final version, shipped until the later loops of Sec 2: same points as v34, plus deadlines at 40/80/90% of the time limit | 0.35592 | 0.038 | 0.8424 |
Visible set, as recorded by each run. Times vary by about 5% between runs of the
same code. The runs' logs (runs/sep26-run1/notes.md,
runs/sep26-run2/notes.md) list every attempt, kept or not, with its measured effect.
FastRiskScore is a heuristic. Unlike RiskSLIM with a full CPLEX licence, it does not prove that its
score is the best one. It is deterministic and single-threaded. It needs numba, which compiles the
search once per machine (about 2 minutes) and caches the result; set
RISKSCORE_NUMBA_CACHE=0 if the package directory is not writable.
Citation
FastRiskScore came out of the autoresearch loop described in Agentic-imodels. If you use either, please cite:
@inproceedings{singh2026agenticimodels,
title={Agentic-imodels: Evolving agentic interpretability tools via autoresearch},
author={Chandan Singh and Yan Shuo Tan and Weijia Xu and Zelalem Gero and Weiwei Yang and Michel Galley and Jianfeng Gao},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems},
year={2026},
url={https://openreview.net/forum?id=EY6O5VNNlV},
}