Skip to content

Learning k-TBNs

pyagrum.ktbn.KTBNDatabaseGenerator samples trajectory CSV databases from a k-TBN, for learning experiments.

class pyagrum.ktbn.KTBNDatabaseGenerator(kdbn)

Section titled “class pyagrum.ktbn.KTBNDatabaseGenerator(kdbn)”

KTBNDatabaseGenerator generates a database of trajectories from a pyagrum.ktbn.KTBN, one CSV file per trajectory. Unlike a plain Bayesian network, a k-TBN describes a stochastic process, so its database is a set of trajectories rather than a flat table of i.i.d. rows: each CSV file has one column per base variable and T data rows, an atemporal variable keeping the same value across all T rows.

Sampling follows the same forward (ancestral) sampling principle as pyAgrum’s Bayesian-network database generator, but never unrolls the k-TBN: it keeps only the k-slice template and slides it forward in time, so memory stays independent of the number of trajectories and their length.

Examples

>>> import pyagrum.ktbn as ktbn
>>> kdbn = ... # a k-TBN
>>> gen = ktbn.KTBNDatabaseGenerator(kdbn)
>>> # 1000 trajectories of length 10 -> out_dir/traj1.csv ... out_dir/traj1000.csv
>>> ll = gen.drawSamples(1000, 10, "out_dir", "traj")

KTBNDatabaseGenerator(kdbn) -> KTBNDatabaseGenerator : Parameters: : - kdbn (pyagrum.ktbn.KTBN) – the k-TBN to sample from (only its k-slice template is copied; the k-TBN itself is not retained)

  • Parameters: kdbn (KTBN)

Generate trajectories, writing one CSV file per trajectory into dirPath. File names are csvBaseName followed by the 1-based trajectory index and .csv (e.g. traj1.csv, traj2.csv, …). Either every trajectory shares the same horizon, or each trajectory gets its own (pass a list of horizons instead of a sample count and a single horizon).

Examples

>>> gen.drawSamples(100, 10, "out_dir", "traj") # 100 trajectories, length 10
>>> gen.drawSamples([8, 10, 12], "out_dir", "traj") # 3 trajectories, own lengths
  • Parameters:
    • nbSamples (int) – the number of trajectories to generate (fixed-length form)
    • nbTimeSlices (int or list *[*int ]) – the horizon shared by every trajectory (fixed-length form), or a list giving each trajectory’s own horizon (variable-length form); every value must be >= k
    • dirPath (str) – directory to write the CSV files into
    • csvBaseName (str) – stem for each file name (index and .csv appended)
    • mode (pyagrum.ktbn.KTBNDatabaseGenerator.VarOrderMode) – the column order of the base variables (default: RANDOM)
    • useLabels (bool) – render values as variable labels rather than modality indices (default True)
    • csvSeparator (str) – the column separator (must not contain a newline)
  • Returns: the log2-likelihood of each generated trajectory
  • Return type: tuple[float, ...]
  • Raises: pyagrum.OperationNotAllowed – if a horizon is smaller than k
  • Returns: the number of base variable columns
  • Return type: int

Set discretized-variable label rendering to the interval label, e.g. "[min,max[".

  • Return type: None

Set discretized-variable label rendering to the (deterministic) interval median.

  • Return type: None

Set discretized-variable label rendering to a uniform random draw within the interval (the default; each labelled export then differs).

  • Return type: None

pyagrum.ktbn.KTBNLearner learns a k-TBN’s structure and/or parameters from trajectory CSVs, at a fixed order k.

KTBNLearner learns a pyagrum.ktbn.KTBN (structure and/or parameters), at a fixed order k, from a set of trajectory CSV files (one file per trajectory, as produced by pyagrum.ktbn.KTBNDatabaseGenerator: one column per base variable, one row per time step).

Internally, k-TBN learning is reduced to three ordinary Bayesian-network learning problems, solved with the same score/algorithm/prior machinery as pyagrum.BNLearner, then glued back together: the transition table (sliding windows of width k) learns the repeating kernel; the initial table (the first k-1 time steps of each trajectory) learns the initial slices; the atemporal table (one row per trajectory) learns the arcs between atemporal variables. Temporal ordering and no-backward-arc constraints are applied automatically.

Examples

>>> import pyagrum.ktbn as ktbn
>>> # atemporal variables inferred from the data
>>> learner = ktbn.KTBNLearner("trajs/", "traj", 500, 2)
>>> learner.useScoreBIC().useGreedyHillClimbing()
>>> model = learner.learnKTBN()
>>>
>>> # or state the classification explicitly (pass an actual `set`, not a list)
>>> learner2 = ktbn.KTBNLearner("trajs/", "traj", 500, 2, {"C", "D"})

KTBNLearner(dirPath, csvBaseName, nbSamples, k, atemporalVars, missingSymbols=[‘?’], induceTypes=True, ignoreMissingSymbols=False) -> KTBNLearner : Structure-learning constructor with the temporal/atemporal classification supplied explicitly.
Parameters: : - dirPath (str) – directory holding the trajectory CSV files - csvBaseName (str) – stem of each file name (1-based index and .csv appended, e.g. "traj" -> traj1.csv, traj2.csv, …) - nbSamples (int) – number of CSV files to read - k (int) – order of the k-TBN; must be >= 2 - atemporalVars (set[str]) – base names of the atemporal (static) variables; every other name found in the CSV header is temporal. Must be an actual Python set (a list/tuple at this position instead selects the atemporal-inferring overload) - missingSymbols (list[str]) – symbols in the CSVs interpreted as missing values - induceTypes (bool) – retype all-numeric columns as integer/range/continuous instead of plain labels - ignoreMissingSymbols (bool) – drop a row carrying a missing symbol instead of refusing the whole database (see nbDroppedRows() for the bias this introduces)

KTBNLearner(dirPath, csvBaseName, nbSamples, k, missingSymbols=[‘?’], induceTypes=True, ignoreMissingSymbols=False) -> KTBNLearner : Same, with the temporal/atemporal classification inferred instead: a base variable is atemporal iff its value never changes across the rows of any single trajectory (a heuristic – prefer the explicit form when the classification is already known).

KTBNLearner(dirPath, csvBaseName, nbSamples, k, bn, atemporalVars=set(), missingSymbols=[‘?’], ignoreMissingSymbols=False) -> KTBNLearner : Variable-schema constructor: types and domains are supplied via a reference pyagrum.BayesNet (one node per base variable, bare names, arcs ignored) instead of inferred from the CSVs. Use this when a variable’s full domain is not guaranteed to appear in the first trajectory – notably atemporal variables, which only ever show one value per trajectory.
Parameters: : - bn (pyagrum.BayesNet) – a BayesNet with one node per base variable, providing the variable types and domains

Forbid an arc from ever appearing in the learned structure.

  • Parameters:
    • tailNode (str) – engine names (e.g. "X[1]", "C") of the two endpoints
    • headNode (str) – engine names (e.g. "X[1]", "C") of the two endpoints
    • tailBase (str) – alternatively, base names of the two endpoints (used together with the slices below)
    • headBase (str) – alternatively, base names of the two endpoints (used together with the slices below)
    • tailSlice (int) – slices of the endpoints (pyagrum.ktbn.KTBN.ATEMPORAL for static)
    • headSlice (int) – slices of the endpoints (pyagrum.ktbn.KTBN.ATEMPORAL for static)
  • Returns: self, for chaining
  • Return type: KTBNLearner

addForbiddenArcAllSlices(tailBase, headBase)

Section titled “addForbiddenArcAllSlices(tailBase, headBase)”

Forbid tailBase -> headBase at every causally-possible slice pair (every lag): tailBase can never be an ancestor of headBase in the learned k-TBN.

  • Parameters:
    • tailBase (str) – base names of the two variables
    • headBase (str) – base names of the two variables
  • Returns: self, for chaining
  • Return type: KTBNLearner

addForbiddenIntraSliceArc(tailBase, headBase)

Section titled “addForbiddenIntraSliceArc(tailBase, headBase)”

Forbid tailBase -> headBase at every intra-slice position (i.e. for every slice t, tailBase[t] -> headBase[t]).

  • Parameters:
    • tailBase (str) – base names of the two temporal variables
    • headBase (str) – base names of the two temporal variables
  • Returns: self, for chaining
  • Return type: KTBNLearner
  • Raises: pyagrum.InvalidArgument – if either endpoint is unknown or atemporal

Force an arc to be part of the learned structure.

  • Parameters:
    • tailNode (str) – engine names of the two endpoints
    • headNode (str) – engine names of the two endpoints
    • tailBase (str) – alternatively, base names of the two endpoints (used together with the slices below)
    • headBase (str) – alternatively, base names of the two endpoints (used together with the slices below)
    • tailSlice (int) – slices of the endpoints (pyagrum.ktbn.KTBN.ATEMPORAL for static)
    • headSlice (int) – slices of the endpoints (pyagrum.ktbn.KTBN.ATEMPORAL for static)
  • Returns: self, for chaining
  • Return type: KTBNLearner

Declare a node as a leaf (forbid it from having any child).

  • Parameters:
    • base (str) – base name of the node (used together with slice)
    • slice (int) – slice of the node (pyagrum.ktbn.KTBN.ATEMPORAL for a static node)
    • name (str) – alternatively, the node’s engine name
  • Returns: self, for chaining
  • Return type: KTBNLearner

Declare a node as a root (forbid it from having any parent).

  • Parameters:
    • base (str) – base name of the node (used together with slice)
    • slice (int) – slice of the node (pyagrum.ktbn.KTBN.ATEMPORAL for a static node)
    • name (str) – alternatively, the node’s engine name (e.g. "X[2]", "C")
  • Returns: self, for chaining
  • Return type: KTBNLearner

Add a candidate edge for MIIC: once at least one edge has been listed, only explicitly listed edges are explored by the structure search.

  • Parameters:
    • tail (str) – engine names of the two endpoints
    • head (str) – engine names of the two endpoints
    • tailBase (str) – alternatively, base names of the two endpoints (used together with the slices below)
    • headBase (str) – alternatively, base names of the two endpoints (used together with the slices below)
    • tailSlice (int) – slices of the endpoints
    • headSlice (int) – slices of the endpoints
  • Returns: self, for chaining
  • Return type: KTBNLearner

Allow or forbid arc additions during structure search.

  • Parameters: allow (bool) – whether to allow arc additions (default True)
  • Returns: self, for chaining
  • Return type: KTBNLearner

Allow or forbid arc deletions during structure search.

  • Parameters: allow (bool) – whether to allow arc deletions (default True)
  • Returns: self, for chaining
  • Return type: KTBNLearner

Allow or forbid arc reversals during structure search.

  • Parameters: allow (bool) – whether to allow arc reversals (default True)
  • Returns: self, for chaining
  • Return type: KTBNLearner
  • Returns: a warning message if the current score and prior are incompatible, an empty string otherwise
  • Return type: str

Copy all score/algorithm/prior/constraint settings from another KTBNLearner (does not copy the database).

  • Parameters: learner (KTBNLearner) – the learner to copy settings from
  • Return type: None
  • Parameters: base (str) – a base variable name (e.g. "X"); an engine name (e.g. "X[1]") is also accepted
  • Returns: the domain size of base
  • Return type: int
  • Returns: domain sizes of the base variables, in the same order as names()
  • Return type: tuple[int, ...]

Undo a previous addForbiddenArc(), same arguments.

eraseForbiddenArcAllSlices(tailBase, headBase)

Section titled “eraseForbiddenArcAllSlices(tailBase, headBase)”

Undo a previous addForbiddenArcAllSlices(), same arguments.

  • Returns: self, for chaining
  • Return type: KTBNLearner
  • Parameters:
    • tailBase (str)
    • headBase (str)

eraseForbiddenIntraSliceArc(tailBase, headBase)

Section titled “eraseForbiddenIntraSliceArc(tailBase, headBase)”

Undo a previous addForbiddenIntraSliceArc(), same arguments.

  • Returns: self, for chaining
  • Return type: KTBNLearner
  • Parameters:
    • tailBase (str)
    • headBase (str)

Undo a previous addMandatoryArc(), same arguments.

Undo a previous addNoChildrenNode(), same arguments.

Undo a previous addNoParentNode(), same arguments.

Undo a previous addPossibleEdge(), same arguments.

  • Returns: always False: incomplete rows are dropped by construction (see nbDroppedRows() to learn whether the CSVs actually had any)
  • Return type: bool
  • Returns: True if the current structure-learning algorithm is constraint-based (MIIC)
  • Return type: bool
  • Returns: True if incomplete rows are dropped rather than rejected outright
  • Return type: bool
  • Returns: True if the current structure-learning algorithm is score-based (BIC, AIC, …)
  • Return type: bool
  • Returns: the order k of the k-TBN being learned
  • Return type: int
  • Returns: (tail, head) engine-name pairs of arcs MIIC flagged as hiding a latent variable (merged from the three internal learners); empty when the algorithm is not MIIC
  • Return type: tuple[tuple[str, str], ...]

Learn the k-TBN’s structure and CPTs.

  • Returns: the learned k-TBN
  • Return type: KTBN

learnParameters(structure, takeIntoAccountScore=True)

Section titled “learnParameters(structure, takeIntoAccountScore=True)”

Learn only the CPTs, using the arc structure of an already-known k-TBN.

  • Parameters:
    • structure (KTBN) – a k-TBN with the same base variables (names and domains) used to construct this learner; a mismatch raises at learn time
    • takeIntoAccountScore (bool) – whether to use the recorded score/prior when estimating parameters (default True)
  • Returns: a new k-TBN with structure’s arcs and freshly learned CPTs
  • Return type: KTBN
  • Returns: base names (no slice suffix), one per base variable, in the original CSV header order
  • Return type: tuple[str, ...]
  • Returns: the number of base variable columns (temporal + atemporal)
  • Return type: int

Number of rows dropped from the internal databases because they carried a missing symbol.

Warning

Dropping rows inflates the log-likelihood computed on the result, and does so more for larger k (a larger k spans more rows per window, so a single missing value costs more of them). Prefer complete trajectories whenever comparing likelihoods across models or orders.

  • Returns: the number of dropped rows, summed across the three internal tables
  • Return type: int
  • Returns: the number of time steps in each trajectory CSV (one entry per sample, in load order); the raw trajectory length, not the transition table’s sliding-window row count
  • Return type: tuple[int, ...]
  • Returns: the number of trajectory CSV files loaded
  • Return type: int

Cap the number of parents of any single node.

  • Parameters: max_indegree (int) – the maximum in-degree
  • Returns: self, for chaining
  • Return type: KTBNLearner
  • Returns: the settings, as (key, value, comment) tuples
  • Return type: tuple[tuple[str, str, str], ...]
  • Returns: a human-readable summary of the learner’s current configuration
  • Return type: str

Use greedy hill-climbing extended with arc-reversal moves.

Use greedy hill-climbing for structure search.

useLocalSearchWithTabuList(tabu_size=100, nb_decrease=2)

Section titled “useLocalSearchWithTabuList(tabu_size=100, nb_decrease=2)”

Use local search with a tabu list.

  • Parameters:
    • tabu_size (int) – the tabu list size (default 100)
    • nb_decrease (int) – the number of non-improving moves tolerated before stopping (default 2)
  • Returns: self, for chaining
  • Return type: KTBNLearner

Use the MDL correction for MIIC’s independence tests.

Use the constraint-based MIIC algorithm for structure search.

Use the NML correction for MIIC’s independence tests.

Disable correction for MIIC’s independence tests.

Use the AIC score for structure learning.

Use the BD score for structure learning.

Use the BDeu score for structure learning.

Use the BIC score for structure learning.

Use the raw log2-likelihood score for structure learning.

Use the MDL score for structure learning.

Use the fNML score for structure learning.

  • Return type: None

Use a Laplace/BDeu-style smoothing prior.

  • Parameters: weight (float) – the prior weight (default 1.0)
  • Returns: self, for chaining
  • Return type: KTBNLearner

pyagrum.ktbn.KTBNAdaptiveLearner additionally selects the order k itself, by exploring every candidate in a range and keeping the best one according to a cross-k model-selection criterion.

class pyagrum.ktbn.KTBNAdaptiveLearner(*args)

Section titled “class pyagrum.ktbn.KTBNAdaptiveLearner(*args)”

KTBNAdaptiveLearner is the order-selecting counterpart of pyagrum.ktbn.KTBNLearner: instead of being given the order k, it explores every candidate k in [kMin, kMax] (kMin starts at 2 and is raised automatically by structural constraints naming a concrete slice), learns one k-TBN per candidate with an internal pyagrum.ktbn.KTBNLearner, and keeps the one with the best cross-k model-selection score (BIC by default).

It exposes the same configuration interface as KTBNLearner (score, algorithm, prior, structural constraints – see that class’s docstrings, valid here too), except settings are only recorded and replayed on each per-k learner at learnKTBN() time; a constraint whose slice does not fit a given candidate is silently skipped for that candidate only.

Examples

>>> import pyagrum.ktbn as ktbn
>>> learner = ktbn.KTBNAdaptiveLearner("trajs/", "traj", 500, kMax=4)
>>> learner.useScoreBIC().useGreedyHillClimbing()
>>> model = learner.learnKTBN()
>>> learner.bestK()
>>> learner.scorePerCandidateK()

KTBNAdaptiveLearner(dirPath, csvBaseName, nbSamples, kMax, atemporalVars, missingSymbols=[‘?’], induceTypes=True) -> KTBNAdaptiveLearner : Parameters: : - dirPath (str) – directory holding the trajectory CSV files - csvBaseName (str) – stem of each file name - nbSamples (int) – number of CSV files to read - kMax (int) – largest order to explore; must be >= 2 - atemporalVars (set[str]) – base names of the atemporal variables (must be an actual Python set) - missingSymbols (list[str]) – symbols in the CSVs interpreted as missing values - induceTypes (bool) – retype all-numeric columns

KTBNAdaptiveLearner(dirPath, csvBaseName, nbSamples, kMax, missingSymbols=[‘?’], induceTypes=True) -> KTBNAdaptiveLearner : Same, with the atemporal classification inferred from the data instead of supplied explicitly (see the equivalent KTBNLearner constructor).

KTBNAdaptiveLearner(dirPath, csvBaseName, nbSamples, kMax, bn, atemporalVars=set(), missingSymbols=[‘?’]) -> KTBNAdaptiveLearner : Variable-schema constructor: types and domains supplied via a reference pyagrum.BayesNet instead of inferred from the CSVs (see the equivalent KTBNLearner constructor).

addForbiddenArcAllSlices(tailBase, headBase)

Section titled “addForbiddenArcAllSlices(tailBase, headBase)”

addForbiddenIntraSliceArc(tailBase, headBase)

Section titled “addForbiddenIntraSliceArc(tailBase, headBase)”

addForbiddenKernelArc(tailBase, lag, headBase)

Section titled “addForbiddenKernelArc(tailBase, lag, headBase)”

Forbid an arc from tailBase, lag slices before the kernel, to headBase in the kernel: tailBase at slice k-1-lag -> headBase at slice k-1, for whichever k is selected. Unlike a plain (base,slice) constraint, the slice moves with the candidate, so it must be expressed relative to the kernel. Adaptive-only: the fixed-k pyagrum.ktbn.KTBNLearner has no moving kernel slice to anchor it to.

  • Parameters:
    • tailBase (str) – base names of the two temporal variables
    • headBase (str) – base names of the two temporal variables
    • lag (int) – the lag before the kernel slice, in [0, kMax)
  • Returns: self, for chaining
  • Return type: KTBNAdaptiveLearner

addMandatoryKernelArc(tailBase, lag, headBase)

Section titled “addMandatoryKernelArc(tailBase, lag, headBase)”

Force an arc from tailBase, lag slices before the kernel, to headBase in the kernel. See addForbiddenKernelArc() for the kernel-relative convention.

  • Parameters:
    • tailBase (str) – base names of the two temporal variables
    • headBase (str) – base names of the two temporal variables
    • lag (int) – the lag before the kernel slice, in [0, kMax)
  • Returns: self, for chaining
  • Return type: KTBNAdaptiveLearner

Data-free check: evaluates the recorded (score, prior) pair as pyagrum.ktbn.KTBNLearner would, so an incompatible combination can be caught before learnKTBN reads any trajectory.

  • Returns: a warning message if the current score and prior are incompatible, an empty string otherwise
  • Return type: str

eraseForbiddenArcAllSlices(tailBase, headBase)

Section titled “eraseForbiddenArcAllSlices(tailBase, headBase)”

eraseForbiddenIntraSliceArc(tailBase, headBase)

Section titled “eraseForbiddenIntraSliceArc(tailBase, headBase)”

eraseForbiddenKernelArc(tailBase, lag, headBase)

Section titled “eraseForbiddenKernelArc(tailBase, lag, headBase)”

Undo a previous addForbiddenKernelArc(), same arguments.

  • Returns: self, for chaining
  • Return type: KTBNAdaptiveLearner
  • Parameters:
    • tailBase (str)
    • lag (int)
    • headBase (str)

eraseMandatoryKernelArc(tailBase, lag, headBase)

Section titled “eraseMandatoryKernelArc(tailBase, lag, headBase)”

Undo a previous addMandatoryKernelArc(), same arguments.

  • Returns: self, for chaining
  • Return type: KTBNAdaptiveLearner
  • Parameters:
    • tailBase (str)
    • lag (int)
    • headBase (str)

Learn and score on the fully observed data only, dropping every row and scoring instance that carries a missing symbol.

Warning

This inflates the log-likelihood, and inflates it more for larger k – so it biases the very order selection this class performs. Inspect scorePerCandidateK() rather than trusting bestK() alone when missing values are frequent.

  • Parameters: ignore (bool) – whether to drop incomplete rows/instances (default True)
  • Returns: self, for chaining
  • Return type: KTBNAdaptiveLearner
  • Returns: True if incomplete rows/instances are dropped (default False)
  • Return type: bool
  • Returns: the largest order explored (the kMax constructor argument)
  • Return type: int
  • Returns: (tail, head) engine-name pairs of arcs the selected model’s MIIC run flagged as hiding a latent variable; empty when the recorded algorithm is not MIIC or none were found
  • Return type: tuple[tuple[str, str], ...]
  • Raises: pyagrum.OperationNotAllowed – if learnKTBN has not run yet

Learn the best k in [kMin, kMax] together with the structure and the CPTs: one internal pyagrum.ktbn.KTBNLearner is built and run per candidate, the recorded configuration is replayed on each, and the k-TBN with the best cross-k score is returned.

  • Returns: the learned k-TBN, at the selected order
  • Return type: KTBN
  • Returns: (k, score) pairs for k = kMin..kMax in ascending k order, from the last learnKTBN call – the values order selection compared to pick bestK() (higher is better; the argmax is bestK)
  • Return type: tuple[tuple[int, float], ...]
  • Raises: pyagrum.OperationNotAllowed – if learnKTBN has not run yet
  • Returns: the recorded configuration, as (key, value, comment) tuples
  • Return type: tuple[tuple[str, str, str], ...]
  • Returns: a human-readable summary of the recorded configuration (candidate order range, algorithm/score/correction/prior, structural constraints), plus the selected k once learnKTBN has run
  • Return type: str

useLocalSearchWithTabuList(tabu_size=100, nb_decrease=2)

Section titled “useLocalSearchWithTabuList(tabu_size=100, nb_decrease=2)”

Select the best k by AIC: a lighter, sample-size-independent complexity penalty than BIC.

Select the best k by BIC (the default): the candidate maximising log2-likelihood minus half the parameter count times log2(sample size).

Select the best k by fNML (factorized Normalized Maximum Likelihood): a data-dependent penalty (unlike BIC/AIC), matching aGrUM’s ScorefNML.

  • Return type: None