8.8.8. ModeSelector#

ModeSelector is a neural network classifier for hadronic FEI \(B\) meson tagging. The FEI assigns a signal probability (sigProb) to each \(B\) candidate independently, using only that candidate’s own decay products. ModeSelector instead uses the full set of FEI \(B\) candidates in the event simultaneously (all candidates with sigProb > 0.001) to predict which physics category the event belongs to and which specific FEI decay mode most likely produced the tag \(B\).

When to use ModeSelector#

ModeSelector builds directly on the FEI: it was trained on FEI skims and requires them as input. The selection on the ModeSelector output then supersedes any additional sigProb cut.

Its main strength is separating \(B^0\) and \(B^+\) events, which substantially reduces crossfeed in analyses that do not fully reconstruct the signal side.

The gain in performance depends on the signal mode: analyses with more inclusive signal sides benefit the most, because the less constrained the signal reconstruction is, the more sensitive it is to crossfeed from the wrong \(B\) type.

Algorithm#

Two-stage network#

ModeSelector runs two neural networks in sequence on each event.

All B candidates -> [feature matrix] -> category network -> B0 / B+ / continuum
                                                 |
                                                 v
                                          main network -> 139-class output -> BplusScore

The category network (3 outputs) classifies the event as \(B^0\), \(B^+\), or continuum based on the full set of candidates.

The main network (139 outputs) then predicts which specific FEI decay mode produced the tag \(B\). The 139 classes are:

  • classes 0-135: signal modes, where the class index (input_id) directly encodes the \(B\) type, decay mode index dmID, and particle/antiparticle sign. Thus the 136 classes are (32 neutral FEI modes + 36 charged FEI modes) \(\times 2\).

  • class 136: bad_tag (combinatorial background)

  • class 137: cross_deltaC1 (\(BB\) crossfeed from the wrong \(B\) charge)

  • class 138: continuum

Each input_id slot can in principle be occupied by multiple FEI candidates (same \(B\) type, decay mode, and charge). When this happens, only the candidate with the highest sigProb is kept per slot (deduplication). The input feature matrix has 9 blocks of per-candidate kinematic and quality variables (FEI signal probability (\(B\) and daughters), vertex fit quality (\(B\) and daughters), D* veto mass differences, deltaE, cosTBTO) across 136 deduplicated slots. To these, 12 event-level variables are appended: 8 event-shape quantities (sphericity, thrust, thrust axis, aplanarity, Fox-Wolfram R2 + 3 harmonic moments) plus the total number of candidates, the input_id of the highest-sigProb candidate, the highest-sigProb input_id from the other \(B\) type, and the experiment number. Sparse columns which do not encode information are excluded from the active network inputs. The main network additionally receives the category network outputs and its predicted \(B\) type as inputs.

D* veto#

The D* veto identifies FEI candidates where the \(B\) was reconstructed with a \(D\) meson directly from the \(B\) decay vertex, but the true decay actually proceeded through a \(D^*\) with a soft pion or \(\pi^0\) that was not reconstructed by the FEI.

These misidentified decays are a significant source of crossfeed: the missing soft pion carries information about the \(B\) flavour and charge, so its absence leads to wrongly identified \(B\) type.

Detection is based on the delta mass difference of a reconstructed \(D^*\) hypothesis:

\[\Delta m_{\mathrm{diff}} = m(D\pi) - m(D) - m^{\mathrm{PDG}}(D^* - D)\]

A small \(|\Delta m_{\mathrm{diff}}|\) indicates a likely \(D^*\) decay. The D* veto produces two variables per \(B\) candidate:

  • Dstp_deltaMassDiff: delta mass difference for the \(D^{*+}\) hypothesis

  • Dst0_deltaMassDiff: delta mass difference for the \(D^{*0}\) hypothesis

These are used as input features and are also stored as candidate ExtraInfo for offline selection.

Outputs#

The following variables are written by the module:

Event-level (EventExtraInfo):

BplusScore

The main signed event-level score. The sign encodes the predicted \(B\) type: positive for \(B^+\), negative for \(B^0\). The magnitude is the maximum main-network probability over candidates present in the predicted sector (background classes excluded). A score close to \(\pm 1\) indicates high confidence; a score near 0 indicates a background-like event.

On events where the predicted sector has no reconstructed candidate, a fallback score is used based on the sum of predicted-sector outputs which encode the predicted \(B\) type.

modeSelector_catBp, modeSelector_catB0, modeSelector_catCont

Category network probabilities for \(B^+\), \(B^0\), and continuum. The predicted sector is \(B^+\) when modeSelector_catBp > modeSelector_catB0.

Candidate-level (ExtraInfo):

modeSelector_eqSigProb

The main-network probability assigned to a candidate’s specific decay mode. Written only on candidates in the predicted sector (\(B^+\) if BplusScore > 0, \(B^0\) if BplusScore < 0). This can serve as a mode-aware complement to sigProb.

modeSelector_rank

Local sector ranking written on deduplicated representative candidates. In the predicted sector, candidates are ranked by modeSelector_eqSigProb; in the non-predicted sector, by sigProb. The two sectors are ranked independently.

Best candidate selection#

Because ModeSelector uses additional information with respect to FEI, its best candidate (rank 1 by modeSelector_eqSigProb in the predicted sector) can differ from the one selected by sigProb alone, though in practice the two agree when the FEI already assigns a clearly dominant signal probability. The mode probability from the main network uses the full event context rather than a single candidate’s features, and its training dataset better reflects the composition expected in data.

modeSelector_rank is written on all deduplicated candidates, in both sectors. The ranking criterion differs by sector:

  • Predicted sector (modeSelector_rank == 1): candidate with the highest modeSelector_eqSigProb.

  • Non-predicted sector (modeSelector_rank == 1): candidate with the highest sigProb. Since modeSelector_eqSigProb is not defined for the non-predicted sector, only sigProb is used for ranking there.

The non-predicted sector contains candidates reconstructed as the \(B\) type opposite to what the category network predicts for the event. These candidates form a crossfeed-enhanced sample that can serve as a calibration or control sample. To restrict the analysis to the predicted sector only, require BplusScore > 0 for \(B^+\) analyses or BplusScore < 0 for \(B^0\) analyses.

How to use#

After running the FEI, apply candidate preselections, build the rest of event and continuum suppression, then add event shape variables and call modeSelector.modeSelector(). The D* veto reconstruction and the neural network evaluation are both set up by this single call.

The preselections Mbc > 5.23, -0.15 < deltaE < 0.1, and cosTBTO < 0.9 were applied during training and are also used in the FEI calibration. They must be reproduced at inference time:

import modularAnalysis as ma
import modeSelector

track_mask = "[[dr < 2] and [abs(dz) < 4] and [pt > 0.2] and [thetaInCDCAcceptance==1]]"
ecl_mask = ("[[[[clusterReg==1] and [E>0.080]] or [[clusterReg==2] and [E > 0.03]] "
            "or [[clusterReg==3] and [E > 0.06]]] and [clusterNHits > 1.5] "
            "and [abs(clusterTiming) < 200] and [thetaInCDCAcceptance==1]]")
roe_mask = ("cleanMask", track_mask, ecl_mask)

for b in ['B+:feiHadronic', 'B0:feiHadronic']:
    # Preselections must match the training setup and FEI calibration
    ma.applyCuts(b, '[Mbc > 5.23] and [-0.15 < deltaE < 0.1]', path=my_path)

    # Build rest of event and continuum suppression (required for cosTBTO)
    ma.buildRestOfEvent(b, path=my_path)
    ma.appendROEMasks(b, [roe_mask], path=my_path)
    ma.buildContinuumSuppression(b, 'cleanMask', path=my_path)

    ma.applyCuts(b, 'cosTBTO < 0.9', path=my_path)

# Some event shape variables are required by ModeSelector
ma.buildEventShape(
    allMoments=False,
    cleoCones=False,
    jets=False,
    collisionAxis=False,
    harmonicMoments=True,
    foxWolfram=True,
    sphericity=True,
    thrust=True,
    path=my_path,
)

# Add ModeSelector (D* veto reconstruction is included automatically)
modeSelector.modeSelector(
    bp_list='B+:feiHadronic',
    b0_list='B0:feiHadronic',
    output_variable='BplusScore',
    path=my_path,
)

# Best candidate selection: keep the rank-1 candidate per list
for b in ['B+:feiHadronic', 'B0:feiHadronic']:
    ma.applyCuts(b, 'extraInfo(modeSelector_rank) == 1', path=my_path)

Models are loaded from the conditions database. The payload names are not given by the user: they are derived from the contract version this release implements, so the training is selected by the globaltag alone. Prepend the recommended ModeSelector performance globaltag in addition to the analysis globaltag:

basf2.conditions.prepend_globaltag('<performance globaltag>')

The payload names can still be given explicitly with payload_cat_model and payload_main_model, which is the way to load the models of an older contract version that this release still supports. Doing so emits a warning, since the training is then no longer selected by the globaltag. A payload built for a newer contract version, or for one that is no longer supported, is rejected with a fatal error rather than used. See Contract version of a weightfile for the scheme.

The following variables are available after running modeSelector(). Event-level outputs are in EventExtraInfo and accessed via eventExtraInfo(...):

  • BplusScore: main output score; positive for \(B^+\), negative for \(B^0\)

  • modeSelector_catBp, modeSelector_catB0, modeSelector_catCont: category network outputs

Candidate-level outputs are in ExtraInfo and accessed via extraInfo(...):

  • modeSelector_eqSigProb: mode-aware (equalized) signal probability (predicted sector only)

  • modeSelector_rank: local sector candidate rank

Training#

Analysts do not normally need to retrain the networks. This section is provided for reference.

Training is performed on run-dependent MC (MCrd) using FEI-skimmed events. FEI calibration factors are applied as per-event weights during training, so the network is conditioned on a sample composition that matches data. The training sample includes both \(BB\) events and continuum, with continuum downweighted to 25% relative importance compared to \(BB\). MC truth matching is applied to assign labels, making use of mostcommonBTagPDG and mostcommonBTagDeltaP.

The category network is trained first. The main network is trained afterwards, receiving both the feature matrix and the category network output as inputs, so the two stages are coupled. The architecture is fully connected with ReLU activations.

All training scripts and configuration are in analysis/scripts/modeSelector/:

  • config.py: training configuration, feature definitions and used FEI calibration weights

  • training/train.py: training pipeline for both networks

  • training/convert_to_onnx.py: export trained models to basf2 MVA weightfiles using ONNX

See analysis/scripts/modeSelector/README.md for full details.

Functions#

modeSelector.modeSelector(bp_list, b0_list, payload_cat_model=None, payload_main_model=None, output_variable='BplusScore', cat_model_path=None, main_model_path=None, addDstarVetoReco=True, training_mode=False, skip_nn_evaluation=False, store_fei_calib_weight=False, debug=False, debug_max_events=10, path=None)[source]#

Add ModeSelector neural network evaluation to a basf2 path.

This function applies a two-stage neural network to FEI B meson candidates to compute an improved signal probability score. The network considers information from all B candidates in the event.

The main output score is stored in EventExtraInfo using output_variable. Auxiliary category and per-candidate mode outputs are stored with the fixed modeSelector_* names.

Parameters:
  • bp_list (str) – B+ meson particle list name to process. Example: ‘B+:feiHadronic’

  • b0_list (str) – B0 meson particle list name to process. Example: ‘B0:feiHadronic’

  • payload_cat_model (str) – Conditions DB payload name for the category model. Used only when cat_model_path is None. None (default) uses the name derived from the contract version this release implements, and the training is chosen by the performance globaltag. Passing a name always emits a warning that the default payload is not used. A payload built for an older contract version runs with that contract’s behaviour if the version is in config.SUPPORTED_CONTRACT_VERSIONS (with a warning); a newer or unsupported version is fatal.

  • payload_main_model (str) – Conditions DB payload name for the main model. Used only when main_model_path is None. Same rules as payload_cat_model.

  • output_variable (str) – Name of the ExtraInfo variable for the output score. Default: ‘BplusScore’

  • cat_model_path (str) – Path to the basf2 MVA weightfile for the category network, as produced by convert_to_onnx.py (a .root file, not a raw .onnx file). If None, loads from conditions database.

  • main_model_path (str) – Path to the basf2 MVA weightfile for the main network, as produced by convert_to_onnx.py (a .root file, not a raw .onnx file). If None, loads from conditions database.

  • skip_nn_evaluation (bool) – If True, skip loading and evaluating the neural networks and fill deterministic placeholder outputs instead. Intended for debugging or timing studies and emits a warning at module initialization.

  • training_mode (bool) – If True, skip NN inference and instead expose per-event features and MC truth as EventExtraInfo/ExtraInfo (modeSelector_feat_XXXX, modeSelector_tr_*, modeSelector_trainSigInputId) for a variablesToNtuple call in the steering script to dump. See analysis/examples/modeSelector/produceTrainingInputs.py. Default: False.

  • addDstarVetoReco (bool) – Whether to add D* veto reconstruction before the NN. Default: True

  • debug (bool) – If True, print the feature vector and network outputs for the first few events. Intended for comparing against a reference implementation. Default: False

  • debug_max_events (int) – Number of events to print when debug is True. Default: 10

  • store_fei_calib_weight (bool) – If True, compute and store modeSelector_feiCalibWeight in EventExtraInfo using the reco path (truth-compatible tag PDG and DeltaP < DELTA_P_THRESH). Returns NaN when reco conditions are not met. Requires mostcommonBTagPDG and mostcommonBTagDeltaP to be defined. Meaningful only on MC. Default: False.

  • path (basf2.Path) – The basf2 path to add the module to.

Notes

Candidate-level ExtraInfo (predicted sector only):

  • modeSelector_eqSigProb: main-network probability for the candidate’s specific decay mode.

  • modeSelector_rank: sector-local rank. Predicted sector ranked by modeSelector_eqSigProb; non-predicted sector ranked by sigProb.

Event-level EventExtraInfo:

  • BplusScore (or the name given by output_variable): signed score, positive for B+ prediction, negative for B0.

  • modeSelector_catB0, modeSelector_catBp, modeSelector_catCont: category network probabilities.

The D* veto reconstruction is added automatically (addDstarVetoReco=True). Pass addDstarVetoReco=False to skip it if already added separately.

modeSelector.addDstarVeto(particleLists, path: Path = None, deltaMassDiffCut: tuple = (-0.02, 0.02), dMassCut: tuple = (-0.03, 0.03), writeExtraInfo: bool = True, skipTreeFit: bool = True)[source]#

Add D* veto reconstruction to the path for B meson particle lists.

This function reconstructs D* candidates by combining D mesons (first daughter of B candidates) with soft pions or pi0s from the Rest of Event, to identify cases where the FEI reconstructed B -> D X but the true decay was B -> D* X.

For B candidates with D0 as first daughter:
  • D*+ -> D0 pi+ (from ROE)

  • D*0 -> D0 pi0 (from ROE)

For B candidates with D+ as first daughter:
  • D*+ -> D+ pi0 (from ROE)

The veto builds its own ROE on private particles that wrap the B candidates, so it neither uses nor creates an ROE related to the B candidates. An ROE built by the user on the input lists, before or after this function, is independent of the veto.

Parameters:
  • particleLists (str or list) – Name(s) of B meson particle list(s) (e.g., ‘B+:feiHadronic’ or [‘B+:feiHadronic’, ‘B0:feiHadronic’])

  • path (basf2.Path) – The basf2 path to add modules to.

  • deltaMassDiffCut (tuple) – Cut on deltaMassDiff (D* mass diff - true mass diff) in GeV.

  • dMassCut (tuple) – Cut on D and D* mass deviation (dM) in GeV.

  • writeExtraInfo (bool) – Whether to write ExtraInfo to particles.

  • skipTreeFit (bool) – If True, skip the vertex TreeFit (significant speedup). The deltaMassDiff will use InvM-based computation instead of fit-based, and chiProb will not be available (stored as NaN). Candidates are ranked by abs(deltaMassDiffInvM) instead of chiProb. Default: True

The following ExtraInfo fields are added to B candidates:

For D0 daughter (D*+ and D*0 veto):
  • Dstp_deltaMassDiff: Delta mass difference for D*+ -> D0 pi+

  • Dstp_chiProb: Vertex fit chi2 probability for D*+ (NaN if skipTreeFit)

  • Dst0_deltaMassDiff: Delta mass difference for D*0 -> D0 pi0

  • Dst0_chiProb: Vertex fit chi2 probability for D*0 (NaN if skipTreeFit)

For D+ daughter (D*+ veto only):
  • Dstp_deltaMassDiff: Delta mass difference for D*+ -> D+ pi0

  • Dstp_chiProb: Vertex fit chi2 probability for D*+ (NaN if skipTreeFit)