FastRiskScore: Fitting sparse integer risk scores via Autoresearch

📄 Paper, 🗂 Doc, 📌 Citation


A risk score is a short list of conditions, each worth a few integer points. You add the points of the conditions that apply and look up the risk for the total in a table. Clinical scores often work this way, so a doctor can compute and check them by hand. Here is an example that predicts whether a breast tumour is benign:

Choosing the conditions and points is a slow integer optimization problem. We used autoresearch to build FastRiskScore, which builds on prior work in the field like RiskSLIM and FasterRisk. On held-out datasets FastRiskScore generally fits much faster than competing baselines, without sacrificing training loss or test AUC.

1. Quickstart & example usage

from imodels import FastRiskScoreClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

model = FastRiskScoreClassifier(k=5).fit(X_train, y_train)
proba = model.predict_proba(X_test)
  • k (default 5): the most conditions the score may use.
  • max_points (default 5): the largest number of points one condition can carry.
  • n_thresholds (default 9): thresholds per numeric column, so the default splits each at its deciles. With more than 19 (for example 99, the percentiles) the search switches to settings tuned for fine thresholds.
  • binarize (default True): numeric columns become x ≤ t conditions and categorical columns one condition per level, so each condition refers to one of the original columns. Set it to False to use binary or small-integer columns as they are.
  • time_limit (default 60 s): the search returns the best score found by then.
The fitted model: points, total score and risk curve

Here is the fitted model, fit in 0.02 seconds:

The fitted model stores the score and its risk curve:

model.points_           # {"worst area <= 906.6": 5, ...}, the lines of the score
model.total_score(X)    # the total points of each row
model.risk(score)       # 1 / (1 + exp(-(model.scale_ * score + model.intercept_)))
model.train_loss_       # the mean training log loss of that risk

The points determine the total score of each row. The risk for a total comes from a logistic curve, risk(s) = 1 / (1 + exp(-(a·s + b))), with a = model.scale_ and b = model.intercept_; both are fit to the training data once the points are chosen. The points are chosen to make the log loss of this curve as small as possible (Sec 3), so the risks in the diagram are calibrated on the training data. RiskSLIM and FasterRisk show the same kind of table; they write b as an integer offset and 1/a as a multiplier.

2. Experimental results

The autoresearch loop. FastRiskScore came out of the autoresearch loop of Agentic-imodels (code on github). Two coding agents (Claude Code with Opus 5.5) started from the same solver, FasterRisk rewritten as a single Python file, and used the same benchmark. One was asked to lower the loss and the other to make the solver faster. In each round an agent changed the solver, ran the benchmark, and kept the change only if the loss went down without the time rising by more than 10%, or the time went down by at least 10% without the loss rising. The agents ran for about 75 hours of autoresearch. The shipped version comes from later loops (below); Fig 1 and Table 1 show it. All solvers are scored the same way. A solver returns integer points in [−5, 5] for at most k of a dataset's binary features. The benchmark then fits the risk curve for those points itself and records 3 numbers: the training log loss of that curve, which is the quantity being minimized; the test AUC of the total score on a held-out 20% split; and the fit time.

Datasets, baselines, and evaluations. The benchmark that the agents optimize over (which we call the "visible set"), has 70 problems: 14 datasets × 5 values of k (3, 4, 5, 7 and 10). Each problem has a 60 s limit and runs on one core. We compare against 19 baselines that run without a commercial solver, from FasterRisk and exact solvers to sparse models with rounding and recent scoring-system packages (listed below).

The 14 visible datasets (from the RiskSLIM and FasterRisk papers, and OpenML)
datasetwhat it isrowsbinary featurespositive rate
adultUCI Adult income, RiskSLIM binarization12,500360.24
bankUCI Bank telemarketing, RiskSLIM binarization12,500570.11
breastcancerUCI Wisconsin breast cancer, 9 cytology scores (1 to 10)68390.35
mammoUCI Mammographic mass, RiskSLIM binarization961140.46
mushroomUCI Mushroom, one-hot8,1241130.48
spambaseUCI Spambase, 57 word and character frequencies (real-valued)4,601570.39
ficoFICO HELOC credit (OpenML 45023)10,0001420.50
compasProPublica COMPAS two-year recidivism (OpenML 42192)5,278320.47
australianStatlog Australian credit690790.45
heartStatlog heart disease270570.44
ionosphereUCI Ionosphere radar returns3512670.36
ilpdIndian liver patient records583790.29
magicMAGIC gamma telescope12,500900.35
habermanHaberman breast cancer survival306270.27

The first six are the files distributed with RiskSLIM, which FasterRisk also uses; breastcancer and spambase keep real-valued columns, and mammo has one small-integer column. The other eight are binarized here: each numeric column becomes indicators x ≤ t at its deciles, each categorical column one indicator per level. Datasets above 12,500 rows are subsampled to 12,500, and every dataset is split 80/20 into training and test rows.

The 27 hidden datasets (from TabArena)
datasetrows / full sizecolumnsbinary featurespositive rate
Amazon_employee_access12,500 / 32,76991070.06
APSFailure12,500 / 76,0001703990.02
Bank_Customer_Churn10,000 / 10,00010540.20
Bioresponse3,751 / 3,7511,7763060.46
blood-transfusion748 / 7484320.24
churn5,000 / 5,000191480.14
coil2000_insurance_policies9,822 / 9,822851760.06
credit-g1,000 / 1,00020870.30
credit_card_clients_default12,500 / 30,000231540.22
customer_satisfaction_in_airline12,500 / 129,880211090.45
Diabetes130US12,500 / 71,518471960.09
E-CommereShippingData10,999 / 10,99910520.40
Fitness_Club1,500 / 1,5006410.30
GiveMeSomeCredit12,500 / 150,00010570.07
hazelnut-contaminant2,400 / 2,400302700.50
HR_Analytics12,500 / 19,15812980.25
in_vehicle_coupon12,500 / 12,684241120.43
Is-this-a-good-customer1,723 / 1,72313750.11
kddcup09_appetency12,500 / 50,0002122940.02
Marketing_Campaign2,240 / 2,240251400.15
NATICUSdroid7,491 / 7,49186800.35
online_shoppers_intention12,330 / 12,33017990.15
polish_companies_bankruptcy5,910 / 5,910643870.07
qsar-biodeg1,054 / 1,054412080.34
seismic-bumps2,584 / 2,58415660.07
taiwanese_bankruptcy6,819 / 6,819943600.03
jm110,885 / 10,885211670.19

Every binary classification dataset of TabArena (OpenML suite tabarena-v0.1) except bank-marketing, heloc and diabetes, which the visible set already contains. They were binarized by the same rule as the visible set, with at most 40 columns kept (those with the highest mutual information with the label on the training split), and never used during development. Builder: evolve_slim/data/build_data.py.

The 19 baselines
baselinefamilysource and how it is run
FasterRiskFasterRiskLiu et al. 2022; package fasterrisk 0.1.10, default settings
FasterRisk, wide searchFasterRiskthe same with a beam of 50 × 50, a pool of 200, 100 swaps and 100 multipliers
RiskSLIMexactUstun & Rudin 2019; risk-slim with the free CPLEX Community Edition
Cutting planes, HiGHSexactRiskSLIM's cutting-plane algorithm reimplemented with the open-source HiGHS solver
SLIMexactUstun & Rudin 2016; its 0-1 loss integer program solved with HiGHS
abess + roundingsparse + roundingZhu et al. 2022; best-subset selection, package abess
fastSparse + roundingsparse + roundingLiu et al. 2022; the L0 path, package fastsparsegams
OKRidge + roundingsparse + roundingLiu et al. 2023; optimal sparse ridge support, vendored
L1 path + roundingsparse + roundingsupports along the L1 logistic path (scikit-learn)
L0Learn + roundingsparse + roundingDedieu et al. 2021; L0L2 logistic path with swaps, R package L0Learn
OKGLM + roundingsparse + roundingLiu et al. 2025; certified k-sparse logistic regression in [−5, 5], vendored, CPU
skscope + roundingsparse + roundingWang et al. 2024; k-sparse logistic regression by splicing, package skscope
Probabilistic scoring listscoring systemHanselle et al. 2025; package scikit-psl, vendored
AutoScorescoring systemXie et al. 2020; reimplemented (R only)
riskscores, annealingscoring systemR package riskscores 1.3.0
riskscores, coordinate descentscoring systemthe same package, its coordinate-descent method
Rounded L1 logisticcommon practiceL1 logistic regression with k features, scaled to a largest point of 5 and rounded
Unit weightingcommon practicethe L1-selected features, each worth +1 or −1 by its sign
imodels SLIMClassifier (before)common practiceimodels 3.0.2 without a commercial solver: a rounded L2 logistic regression

The baselines came from two searches for methods that can be run without a commercial solver: one of packages for risk scores and sparse logistic regression, and a later one of 2023 to 2026 papers that cite FasterRisk, RiskSLIM, fastSparse, OKRidge, scoring lists or the certifier of Liu et al. (ICML 2025). Each "+ rounding" method gives a sparse real-valued solution with at most k features, which FasterRisk's rounding step turns into points. Penalized methods (riskscores, L0Learn, the L1 path, imodels' SLIMClassifier) have their penalty searched until at most k points are nonzero, and points are clipped to [−5, 5]. scikit-psl and okridge pin numpy < 2 and are vendored with small fixes (a numpy 2 incompatibility in scikit-psl, a call to a removed method in okridge), and scikit-psl gains a stop after k stages. AutoScore is only available in R, so its recipe (rank by random-forest importance, fit a logistic regression, divide by the smallest coefficient and round) is reimplemented on the binary features. RiskSLIM has no open-solver version, so its cutting-plane algorithm is reimplemented with HiGHS, re-solving the integer program each round and starting from rounded heuristic solutions as RiskSLIM does; it proves optimality on 6 of the 70 visible problems. L0Learn runs from its R package, since the Python package has no build past Python 3.10, and the R methods run in one persistent R process per worker so that R's start-up is not timed. OKGLM has no intercept, so a constant column is added. imodels' previous SLIMClassifier solves its integer program only with a commercial solver installed; without one, which is the default, it rounds an L2 logistic regression. Probabilistic scoring lists are killed at 3× the time limit on 10 visible and 51 hidden problems, and riskscores often passes the time limit, since it checks the clock only between fits; a solver without a score is charged the base rate. Left out: scorepyo (abandoned, pinned to 2022 packages), weight-of-evidence scorecards (optbinning, scorecardpy; not sparse integer points), GroupFasterRisk (the same as FasterRisk without feature groups), recent methods without public code (a GPU branch and bound for k-sparse GLMs, net-benefit risk scores) and those that optimise a different objective (boosted risk scores, SHAPoint, AgentScore).

FastRiskScore is fast and accurate. Fig 1 plots time per problem against the training loss (Fig 1A-C) and the test error (Fig 1D). In all cases, FastRiskScore achieves the best performance (lower left). The loss is shown as the excess over the lowest loss any solver reached on each problem, averaged over problems. A solver that returns no score on a problem is charged the loss of predicting the base rate there.

  • Autoresearch run 1 (lower loss)
  • Autoresearch run 2 (faster)
  • Discarded attempts
  • Baselines
  • FastRiskScore

Visible set (14 datasets × 5 values of k, 60 s per problem)

Hidden set (27 datasets × 5 values of k, 60 s per problem)

Hidden set, full size (k = 5, 10 min per problem)

Hidden set, test error (27 datasets × 5 values of k, 60 s per problem)

Fig 1. Training loss and test error against fit time. FastRiskScore (star) is about 180× faster than FasterRisk on the visible set and 225× on the held-out sets (by the median), and has a lower loss on 55 of 70, 94 of 135 and 21 of 27 problems (a higher one on 0, 8 and 0), and has the lowest test error of the integer scores. The closest baselines are sparse models rounded with FasterRisk's rounding step. A: the visible set that the loop was scored on, with every version either run tried (discarded attempts are faded). B: the same protocol on 27 held-out TabArena datasets, for the versions each run kept. C: the held-out datasets at full size, up to 120,000 training rows, at k = 5 with 10 minutes per problem. D: the test error (1 − AUC) of the solvers in B, averaged over the 135 held-out problems; the axis stops at 0.275, so the solvers above it (RiskSLIM 0.432, SLIM 0.358, cutting planes 0.333 and scoring lists 0.319) are not shown. Each point is one solver; the x-axis is its geometric-mean fit time, and in A to C the y-axis is its mean excess training loss over the best loss any solver reached on each problem (log scale). Large points are the mean of three runs. In C no other solver shown beats FastRiskScore on any dataset, so its star sits at the bottom of the axis. All runs used one core per problem on a Linux server (two Xeon Platinum 8468V), with 14 problems at a time (27 in C).

Extra details on experimental settings

Scoring. The benchmark rejects any answer whose points are not integers, fall outside [−5, 5] or use more than k features; no solver in the figure had such an answer. Given the points, the risk curve is fit by Newton's method on the two numbers a and b, a convex problem. Each solver process first fits a tiny warm-up problem, untimed, so numba compilation is not counted; FastRiskScore's first fit on a new machine takes about 2 minutes longer for that reason.

Best known losses. The excess loss is measured against the lowest loss any integer solver reached on the problem in any of these runs, including FasterRisk with every search width raised (a beam of 50 × 50, a pool of 200 and 100 multipliers) and a 20-minute limit, which takes about 12× its default time.

RiskSLIM. The free CPLEX Community Edition refuses problems above 1,000 variables or constraints, and RiskSLIM's cutting planes pass that limit on wider datasets; it returned no score on 25 of 70 visible problems and 102 of 135 hidden ones, and is charged the base rate there. With a full CPLEX licence it would do better, but within these time limits it still certifies optimality only on small problems, as the FasterRisk paper also found. RiskSLIM and SLIM run until their time limit on most problems, so their times sit near 60 s.

C. One run per solver at k = 5 on each of the 27 datasets at full size, 27 at a time, with a 10-minute limit (killed at 30 minutes). RiskSLIM stops at the CPLEX Community Edition's size limit on 22 datasets; SLIM passes its limit on 25, since HiGHS finishes its current step before stopping, and on the remaining 2 returns a proven optimum of its 0-1 loss.

Table 1: loss, test AUC and time of every solver on the three suites

Table 1. Mean training log loss, mean test AUC and geometric-mean fit time (s) of every solver on the three suites of Fig 1. Eight of the baselines (cutting planes, abess, fastSparse, OKRidge, the L1 path, scoring lists, AutoScore and unit weighting) were added after a search for baselines that can be run, and five more (L0Learn, OKGLM, skscope and two versions of riskscores) after a search of papers that cite them (details). The last row is not a risk score: it is FasterRisk's real-valued solution before rounding, shown as a reference.

Visible (70 problems)Hidden (135 problems)Hidden, full size (27)
solverlossAUCtimelossAUCtimelossAUCtime
FastRiskScore (ours)0.35580.8420.0110.32690.7940.0150.32790.7920.020
FasterRisk0.36030.8411.510.32720.7933.280.32820.7904.82
FasterRisk, wide search0.35820.84118.20.32710.79347.3–––
RiskSLIM (CPLEX CE)0.45240.71813.60.40460.5689.980.40310.56016.2
Cutting planes, HiGHS (RiskSLIM's method)0.40970.776570.37390.667600.36020.695600
SLIM (HiGHS)0.43710.76749.30.38950.64260.20.38350.647542
abess + sequential rounding0.37310.8420.2070.33030.7890.5120.33170.7880.790
fastSparse + sequential rounding0.37380.8360.1230.33030.7910.2070.33040.7900.304
OKRidge + sequential rounding0.36540.84120.40.32860.79026.50.32970.786214
L1 path + sequential rounding0.40870.8130.0990.35460.7440.2310.35810.7440.370
Probabilistic scoring list0.41850.78713.80.36230.68151.20.33660.763127
L0Learn + sequential rounding0.36330.8430.8800.32900.7922.840.32930.7904.21
OKGLM (certified k-sparse) + sequential rounding0.36090.84225.30.32930.79231.40.32960.793175
skscope + sequential rounding0.37510.8370.2220.33180.7910.7200.33250.7921.12
riskscores (coordinate descent)0.44760.75712.70.38170.66636.60.37130.699111
riskscores (annealing)0.48760.70138.30.40970.58570.10.38210.665283
AutoScore0.40000.8210.2600.34610.7630.4950.34690.7680.869
Rounded L1 logistic0.40960.8120.0680.35470.7430.1800.35800.7440.292
Unit weighting0.43110.8040.0620.36650.7320.1800.36650.7390.293
imodels SLIMClassifier (before)0.42600.7890.2580.35830.7360.8100.36120.7321.42
Real-valued k-sparse (not integer)0.35740.8440.8670.32650.7951.560.32750.7942.27
Table 2: six later loops on versions of both suites with 99 thresholds per numeric column instead of 9

The suites above split each numeric column at its 9 deciles. We also built versions of the visible and hidden suites that split each numeric column at its 99 percentiles, which gives the search many more features to choose from (990 instead of 90 on magic). The six RiskSLIM datasets were rebuilt from their raw OpenML sources for this. Six more loops ran on the fine visible suite with the same keep rule. Loops 1 and 2 both started from FastRiskScore; each later loop started from the best version of the one before. From loop 3 on, a change was also rejected if it lowered the mean test AUC by more than 0.001. The fine hidden suite (27 datasets × 5 values of k) was not used during the loops. Table 2 gives the best version of each loop that kept a change, on both.

Table 2. Mean training log loss, mean test AUC and geometric-mean fit time (s) on the fine suites (99 thresholds per numeric column). Run 2's version is the first release of FastRiskScore; each loop row is that loop's best version and its main change; the last row is the shipped solver with n_thresholds > 19.

Fine visible (70)Fine hidden (135)
solverlossAUCtimelossAUCtime
FasterRisk0.342820.8482.160.324480.7924.62
Run 2's version (first release)0.341190.8500.0680.324160.7920.153
Loop 1: wider local search with kicks0.338820.8470.2590.323930.7930.322
Loop 2: thresholds of one column handled together0.339820.8480.0320.324000.7930.050
Loop 3: minimum support per feature0.339980.8510.0230.324160.7930.042
Loop 4: faster0.339960.8510.0110.324140.7930.021
Loop 5: exact polish of the best score0.339760.8510.0110.324140.7930.021
FastRiskScore (shipped, fine settings)0.340790.8520.0110.324140.7940.020

Versions: loop 1 f31_fastsame, loop 2 g24_fastcol, loop 3 h8_fastbeam, loop 4 i30_final, loop 5 j18_poliso. Hidden-suite AUCs to four decimals: FasterRisk 0.7917, run 2's version 0.7924, loops 1 and 2 0.7927, loop 3 0.7930, loops 4 and 5 0.7932, shipped 0.7937. Times vary by about 5% between runs; the loop 1 and 2 hidden rows were run while another loop used the same machine.

  • Loss. Every loop beats FastRiskScore on the fine visible suite, but on the fine hidden suite the gains are small: 0.0002 at most. Loop 1 has the lowest loss on both suites, at 6 to 15 times the hidden-suite time of the later loops.
  • Feature handling. Loop 2 found that the 99 indicators of a numeric column are nested (x ≤ t for increasing t), and handled each column's indicators as one ordered variable. This made the search faster and let the beam pick the best threshold of a column instead of filling with neighbouring thresholds of one column.
  • AUC. Loop 3 required each chosen indicator to cover at least √N training rows on each side. Without it, the search picks thresholds that fit a handful of training rows. Against loop 2, it raised visible AUC by 0.003 for a loss increase of 0.0002, and the AUC gain held on the hidden suite (0.7927 to 0.7930).
  • Speed. Loop 4 reached about the same loss as loop 3 in half the time: 0.021 s on the hidden suite, 7× faster than FastRiskScore and 220× faster than FasterRisk.
  • Saturation. Loop 5's only kept change, an exact polish of the final score, lowered the visible loss by 0.0002. Almost all of that came from two near-separable problems (ionosphere at k = 7, mushroom at k = 10), and it changed 2 of the 135 hidden problems. On the problems where the polish runs, no single change of one feature's points, and no swap of one feature, lowers the loss of the final score. A sixth loop (30 attempts: new starting points, tabu walks, pair moves) kept nothing; running its whole search 4 times longer lowered the mean loss by 0.00001.
Table 3: how the shipped version was found (two loops on the decile suites), compared with run 2's version

The fine-track versions are 3× faster than run 2's version, but on the decile suites they lost 0.0014 in visible loss, mostly on spambase and ionosphere. Two causes accounted for most of it: loop 3's minimum support per feature, and beam rules that kept one threshold per variable and one child per set of variables, which removed good feature sets on decile data. A loop on the decile visible suite started from loop 5's version with the minimum support switched off and kept four changes: candidates whose real-valued coefficient rounds to 0 are ranked last in the beam (spambase, whose large-scale columns collapsed 3 of 5 final feature sets to 2 features), the search kernels are compiled before the first fit (about 4 s had been spent compiling inside a few fits), the beam diversity rules are switched off, and the beam is trimmed back to 8 parents. A second loop from that version found nothing that held on the hidden suite. The package ships the second loop's final version (n17_lean, with the same hidden-suite results), and switches to the fine-track settings when n_thresholds > 19. Table 3 compares it with run 2's version.

Table 3. Mean training log loss, mean test AUC and geometric-mean fit time (s) of the shipped version and run 2's version on the suites of Table 1.

Visible (70 problems)Hidden (135 problems)Hidden, full size (27)
solverlossAUCtimelossAUCtimelossAUCtime
FasterRisk0.360290.8411.490.327200.7933.270.328230.7904.82
Run 2's version (first release)0.355920.8420.0350.326950.7940.0570.327890.7930.079
FastRiskScore (shipped)0.355830.8420.0110.326880.7940.0150.327890.7920.020

AUCs to four decimals, visible / hidden / full size: run 2's version 0.8424 / 0.7937 / 0.7926, shipped 0.8421 / 0.7939 / 0.7920. Times vary by about 10% between runs.

On the hidden suite the shipped version has a lower loss than run 2's version on 23 of the 135 problems and a higher one on 10, and it is 3 to 4× faster on the three suites. Test AUC is within 0.0006 of run 2's version on every suite.

3. Methods: FastRiskScore

The problem is the one RiskSLIM and FasterRisk solve. Over integer points w with at most k nonzero entries, each in [−5, 5], minimise

\[ \min_{w}\;\; \min_{a,\,b\,\in\,\mathbb{R}}\; \underbrace{\frac{1}{n}\sum_{i=1}^{n} \log\!\left(1 + e^{-\tilde y_i\,(a\, x_i^\top w \,+\, b)}\right)}_{\text{log loss of the risk curve for the total score } x_i^\top w} \]

with \(\tilde y_i \in \{-1, +1\}\). The inner minimum over a and b is the risk curve of Sec 1. FasterRisk first finds good real-valued coefficients and then rounds them, trying a few multipliers 1/a with the intercept fixed by the rounding. FastRiskScore keeps FasterRisk's first stage, rewritten to run faster, and replaces the rounding stage with a search over the integer points, where each candidate is scored by the loss above with a and b fit again. The shipped version has five steps:

  1. Compress the data. Identical rows (same features and label) are merged into one row with a count, and the nested thresholds of one numeric column (x ≤ t for increasing t) become one ordered variable, so the counts for every threshold come from one table.
  2. Beam search for real-valued solutions (FasterRisk's): grow the set of features one at a time from the 8 best sets. Each set proposes its 10 most promising features, ranked by a one-step Newton bound, at most one per threshold band of a variable, and the 20 best proposals are fit by Newton's method on rows grouped by their values on the chosen features. A set whose coefficient on a raw numeric column would round to 0 is ranked last.
  3. Round each of the 8 best final solutions at 20 scales, and at 4 more where the largest points are clipped at the bound so the others keep finer ratios, keeping the rounding with the lowest loss after refitting a and b.
  4. Local search on the integer points from the best roundings: change one feature's points or add a feature, move a threshold along its column, then swap one feature for another; finally try the best swaps followed by a real-valued refit and rounding. The loss of a move is computed from a table of (score, feature value) counts; an estimate ranks the moves and the best are checked exactly.
  5. Exact polish when the score nearly separates the classes: every single change is checked exactly, skipping candidates whose lower bound cannot beat the current loss.

With more than 19 thresholds per column the search also requires each chosen feature to cover at least √n rows on each side, keeps one threshold per variable in the beam, and uses 10 sets in the beam and 5 roundings; these settings come from the fine-threshold loops of Sec 2.

Loop findings

Searching the integer points directly, scored by the refit loss (both runs)

Both runs found this change first, and it lowered the loss more than any other: on the visible set it moved the mean excess loss from about +0.002 to −0.001 (measured against the original best known losses) in each run. FasterRisk judges a rounding by its loss at a fixed multiplier, but a risk score is used with its fitted curve, so a rounding that is worse at the fixed multiplier can be better once a and b are fit again. On spambase at k = 3, FasterRisk's rounding sets one of its three features to zero, because at that feature's scale the best integer point rounds to 0. The local search puts the feature back and lowers the loss by 0.095.

At first the local search added 60 to 80% to the run time, because every candidate move required fitting a and b again over all rows. Two changes made it cheap. The loss is computed from the counts of positive and negative rows at each value of the total score (with 5 features the total takes at most 51 values), and moves are ranked by an estimate so that only the best two are checked exactly. Run 2 found that its first estimate was too optimistic for moves that change the spread of the scores (0.58 estimated against 0.63 true on compas at k = 3). It fixed this by recomputing the cells a move touches after one Newton step in (a, b).

Speed changes

A profile of FasterRisk on bank at k = 10 (19.9 s) showed 82% of the time in coordinate descent, which reads every row for each coordinate step. Merging identical rows and fitting by Newton's method on the few chosen features made both runs 5× faster. Run 2 then grouped each beam candidate's rows by their values on its features (another 2.5×), moved the creation of each new candidate into one compiled function together with smaller changes (−34% time), and computed the gradient and swap screens as single matrix products (−14%, then −12%). After these changes each step reads the data once instead of once per candidate.

The integer search also made parts of FasterRisk's first stage unnecessary. Run 2 removed FasterRisk's diverse pool (swapping each feature for up to 50 others) and started the local search from the 10 beam solutions instead, which gave a lower loss. Run 1 kept the pool but reduced it from 50 to 10 solutions and the rounding multipliers from 20 to 5, with no change in loss.

Changes that didn't yield improvements
  • Local search from an empty score, without the beam search: much worse (excess loss +0.010 against −0.001), so the features chosen by the beam search matter.
  • Cheaper ways to rank beam candidates before fitting them (a one-step score test, a one-coordinate fit): all raised the loss by 0.0002 to 0.0007.
  • Smaller beams: 5 parents instead of 10 raised the loss by 0.003.
  • Random restarts and perturbations of the local search: large gains on a few problems where it got stuck (mushroom at k = 7), but less than 0.0001 on average, for 2 to 3 times the time. 300 random restarts per problem on 15 problems beat the final version on only 4.
  • More starts, deeper swaps, rescaling all points at once, refitting the best set of features and rounding again at several resolutions: gains of at most 0.0002, for 20 to 60% more time, so none passed the keep rule.
Full table of kept versions
versionwhat changedmean losstime (s)test AUC
Run 1 (lower loss)
pyfasterrisk_v1FasterRisk in one file: beam search, diverse pool, star-ray search with sequential rounding0.360291.4440.8406
v3_fastnewtonthe same algorithm on unique (x, y) rows with counts, numba kernels, Newton on the support in place of coordinate descent0.359200.2790.8415
v5_tunetopbeam: fine-tune only the children with the best one-coordinate fit0.359240.2110.8424
v6_numbaraystar-ray search and sequential rounding in one numba kernel0.359240.1760.8424
v9_pool10integer local search scored by the calibrated loss, on score histograms; rounding pool 50 → 100.356610.1630.8427
v11_pooltunediverse pool: fine-tune only the most promising swaps0.356610.1430.8427
v24_child5rescale and paired ±1 moves to polish the best result; beam child size 10 → 50.356310.1420.8426
v27_cheap1dranking fits capped at 3 Newton steps; 20 swap attempts per feature0.356320.1220.8425
v28_ray55 multipliers in the star-ray search instead of 200.356320.1050.8425
v29_tune1row deduplication by a random projection; fewer fine-tuned candidates0.356360.0870.8427
v34_2startslocal search in numba on the support; 2 starts instead of 30.356400.0740.8431
v36_robustend of run: same points, plus a deadline and kernels compiled at import0.356400.0710.8431
Run 2 (faster)
v2_numbaFasterRisk in numba on unique (x, y) rows; projected Newton in place of coordinate descent0.359810.2960.8414
v5_nopoolinteger local search scored by the calibrated loss (value changes, additions, swaps); diverse pool dropped0.356370.2640.8427
v7_groupsbeam nodes keep rows grouped by (y, support pattern): at most a few hundred groups instead of all rows0.356370.1070.8427
v14_numbachildfaster data setup, swap screen, score binning by counting, visited-state memory; child creation in one numba call0.356390.0710.8428
v16_swaplastvalue changes and additions before swaps; swaps only when those fail0.356220.0630.8427
v17_calroundrounding by the calibrated loss over 20 scales, replacing star-ray sequential rounding0.356020.0660.8430
v18_roundnbthat rounding in one numba kernel0.355970.0570.8430
v24_blasscreenbetter move estimate (touched cells exact after a Newton step in (a, b)); swap screen as two BLAS products0.355920.0490.8428
v30_lazycodescolumn codes computed lazily; exact checks skipped when the estimate rules a move out0.355920.0420.8428
v34_gemmswap screens for all k removals, and beam gradients for all parents, as one matrix product each0.355920.0370.8424
v35_scratchrun 2's final version, shipped until the later loops of Sec 2: same points as v34, plus deadlines at 40/80/90% of the time limit0.355920.0380.8424

Visible set, as recorded by each run. Times vary by about 5% between runs of the same code. The runs' logs (runs/sep26-run1/notes.md, runs/sep26-run2/notes.md) list every attempt, kept or not, with its measured effect.

FastRiskScore is a heuristic. Unlike RiskSLIM with a full CPLEX licence, it does not prove that its score is the best one. It is deterministic and single-threaded. It needs numba, which compiles the search once per machine (about 2 minutes) and caches the result; set RISKSCORE_NUMBA_CACHE=0 if the package directory is not writable.

Citation

FastRiskScore came out of the autoresearch loop described in Agentic-imodels. If you use either, please cite:

@inproceedings{singh2026agenticimodels,
      title={Agentic-imodels: Evolving agentic interpretability tools via autoresearch},
      author={Chandan Singh and Yan Shuo Tan and Weijia Xu and Zelalem Gero and Weiwei Yang and Michel Galley and Jianfeng Gao},
      booktitle={The Fortieth Annual Conference on Neural Information Processing Systems},
      year={2026},
      url={https://openreview.net/forum?id=EY6O5VNNlV},
}