Rules and Checks Reference
tff runs two categories of quality guardrails: Architectural Checks and Linter Rules. All of these are configured in the fitness_functions.yaml file in the root of your project.
🛠️ Auto-Fixer (--fix)
tff includes a built-in auto-fixer that can automatically resolve simple violations. By running tff lint --fix, tff will modify your source files to fix the following issues:
- No Positional GROUP BY/ORDER BY (
nopositionalgroupby,nopositionalorderby,nopositionalgroupbyororderby): Rewrites integer positional references inGROUP BYandORDER BYclauses to explicit column names or select aliases using AST modification. - Nested Subqueries in Final SELECT (
sqlcomplexity): Refactors inline subqueries inFROM (SELECT ...) aliasandJOIN (SELECT ...) aliasclauses of the finalSELECTstatement into named Common Table Expressions (WITH alias AS (...)). - Metadata (
nomissingowner,nomissingdescription):- dbt: Automatically appends or scaffolds
schema.ymlmetadata configs with"TODO: Add owner"and"TODO: Add description"templates. - SQLMesh: Inline-updates the
MODELblock in the model.sqlfile to addowneranddescriptionheaders. - Dataform: Injects
description: "TODO: Add description"andbigquery: { labels: { owner: "TODO: Add owner" } }inside.sqlxconfig { ... }blocks, or scaffolds a minimalconfig { ... }header if absent.
- dbt: Automatically appends or scaffolds
⚡ Performance: Parallel Traversal & AST Caching
For enterprise DAGs consisting of hundreds or thousands of transformation models, tff provides multi-core parallelism and persistent disk-based caching:
- Parallel Model Loading & Parsing: AST parsing is dispatched across a pool of worker processes (
ProcessPoolExecutor) during model loading. - Parallel Duplicate CTE Fingerprinting: CTE extraction, AST normalization, and cryptographic hashing run concurrently across models.
- Persistent AST Caching: Precomputed ASTs are cached in
.tff_cache/astkeyed by SQLGlot version, target dialect, and model SQL SHA-256 hash. Repeat evaluations achieve sub-second execution speeds. - Threaded Rule Execution: Model-level rules and check definitions run concurrently via thread pools (
run_parallel_model_rule). Models are dynamically batched to minimize worker thread scheduling overhead.
Configuration in fitness_functions.yaml
# Root configuration options:
provider: dbt # Optional: explicitly set pipeline provider (dbt, sqlmesh, dataform)
workers: 4 # Optional: number of worker processes (default: auto, capped at CPU count)
cache_ast: true # Optional: toggle persistent AST caching (default: true)
cache_dir: ".tff_cache" # Optional: persistent cache directory (default: ".tff_cache")
You can also override these on the CLI via --workers <N>, --no-cache, and --clear-cache, or via the TFF_WORKERS environment variable.
Task Batch Tuning (TFF_CHUNK_SIZE)
When evaluating model-level rules across large repositories (>1,000 models), creating individual tasks per model can introduce thread pool scheduling and synchronization overhead. tff automatically groups eligible models into task batches:
-
Dynamic Default Formula: When
This creates approximately 4 batches per worker thread to ensure even thread load balancing, bounded between a minimum of 1 and a maximum of 100 models per batch.TFF_CHUNK_SIZEis unset, the batch chunk size is calculated dynamically: -
Environment Variable Override: You can set
TFF_CHUNK_SIZEin CI or local environments to tune batch sizes for your workload:Setting# Tune batch size for high-volume repositories in CI or local runs export TFF_CHUNK_SIZE=50 tff lintTFF_CHUNK_SIZE=1reverts to single-model execution per task.
Shared Layer Filtering Configuration
Most checks and rules inherit a common layer filtering schema. This allows you to apply guardrails selectively based on the pipeline layer a model belongs to:
rules:
some_rule:
enabled: true # Toggle the rule on or off (default: true)
skip_layers: [staging] # List of layers where this rule should NOT run
only_layers: [marts] # If specified, the rule ONLY runs on these layers
1. Architectural Checks
Architectural checks evaluate the structure, dependencies, and layout of your entire project DAG. They are run via the tff lint or tff health CLI commands.
Layer Integrity (layer_integrity)
- What it checks:
- Unidirectional Dependency Flow: Ensures models in upstream layers do not depend on models in downstream layers (as defined by the index order in
layers.order). - Mart Domain Isolation: Ensures models within the
martslayer (models/marts) do not depend on models in other domains within themartslayer (e.g.,models/marts/financecannot depend onmodels/marts/marketing). - How to configure:
Defined under
checks.layer_integrityinfitness_functions.yaml.Layers are located at the first level within the models/ directory, e.g.layers: order: [staging, core, marts] # Bottom-to-top hierarchy order checks: layer_integrity: enabled: trueorder: [staging, core, marts]assumesmodels/staging,models/core, andmodels/martsdirectories. These can contain subdirectories with models.
Custom Exclusions (custom_exclusions)
- What it checks:
- Enforces custom dependency boundaries. It blocks defined layer/domain dependencies and supports specifying whitelist exceptions. Exclusions can be defined directly in
fitness_functions.yamlor in a separate JSON file. - How to configure:
Defined directly under
exclusionsandallowed_exceptions(or underchecks.custom_exclusions) infitness_functions.yaml:Alternatively, you can point to an external JSON exclusions file (e.g.exclusions: - source_layer: core target_layer: derived - source_layer: core source_domain: finance target_layer: marts target_domain: marketing allowed_exceptions: - model: derived.model_name dependency: core.dependency_name checks: custom_exclusions: enabled: truelinter_exclusions.json):The exclusions file (e.g.,exclusions_path: linter_exclusions.json # Relative to project root checks: custom_exclusions: enabled: truelinter_exclusions.json) has the following structure: exclusions: A list of blocked dependencies. If a model in thetarget_layer/target_domaindepends on a model in thesource_layer/source_domain(which is the source of the dependency relation), a violation is raised. Omitting domain fields matches all domains in that layer.allowed_exceptions: Specificmodel$\rightarrow$dependencypairs to allow even if they match an exclusion rule.
#### Metadata/Tag-driven Fallback
If models are organized by functional theme rather than layer directories, layer_integrity and custom_exclusions will automatically fall back to using tags and metadata:
* Layer: Set a tag matching one of the layers in the layers.order list, or define a layer key in model metadata.
* Domain: Prefix a tag with domain:, e.g., domain:finance, or define a domain key in model metadata.
Example (Functional Theme Layout):
Assume a model is located at models/finance/payments_cleared.sql. Since finance is not in your configured layers.order, tff's directory parser cannot determine the layer automatically. You can explicitly tag/annotate it:
-
dbt (
schema.yml): -
SQLMesh (
payments_cleared.sql):
#### YAML-based Configuration
Instead of or in addition to a JSON file, custom exclusion rules and exceptions can also be defined directly in fitness_functions.yaml under checks.custom_exclusions using tags and metadata selectors:
checks:
custom_exclusions:
enabled: true
exclusions:
# Exclude public models from depending on pii models via tags
- source_tag: "pii"
target_tag: "public"
# Exclude based on metadata key-value pairs
- source_meta:
team: "marketing"
target_meta:
team: "finance"
# Combine layer/domain with tag/meta selectors
- source_layer: "core"
source_tag: "confidential"
target_layer: "marts"
allowed_exceptions:
- model: "marts.public_model"
dependency: "core.confidential_model"
Schema Contracts (schema_contracts)
- What it checks:
- Enforces schema structural parity between related models to ensure they stay in sync. Contracts can be configured directly in
fitness_functions.yamlor in an external JSON file. - How to configure:
Defined under
contract_groups(or underchecks.schema_contracts) infitness_functions.yaml:Alternatively, you can point to an external JSON contract groups file (e.g.contract_groups: column_parity_groups: - reference: models/core/dim_customer_ref.sql exclude_columns: [created_at, updated_at] members: - models/core/dim_customer_replica.sql dimension_parity_groups: - left: models/core/fact_sales.sql right: models/core/fact_orders.sql checks: schema_contracts: enabled: truelinter_contract_groups.json):The schema contracts file (e.g.,contract_groups_path: linter_contract_groups.json # Relative to project root checks: schema_contracts: enabled: truelinter_contract_groups.json) supports two contract formats:
#### 1. Column Parity Groups Enforces that member models contain the exact same columns in the exact same order as a reference model.
{
"column_parity_groups": [
{
"models_dir": "models/core",
"reference": "dim_customer_ref.sql",
"exclude_columns": ["created_at", "updated_at"],
"reference_substitutions": {
"customer_id": "id"
},
"members": [
{
"file": "dim_customer_replica.sql",
"substitutions": {
"cust_id": "id"
}
}
]
}
]
}
models_dir: The base directory within the project root for the files.
* reference: The SQL file of the source-of-truth model.
* exclude_columns (optional): Columns to ignore in the comparison.
* reference_substitutions (optional): Maps reference columns to a common name for comparison.
* members: The list of member models. Each member can define own substitutions to align column names.
#### 2. Dimension Parity Groups Enforces that two models contain the exact same set of dimension columns, regardless of their select order.
{
"dimension_parity_groups": [
{
"models_dir": "models/core",
"left": {
"file": "fact_sales.sql",
"exclude_columns": ["revenue"]
},
"right": {
"file": "fact_orders.sql",
"exclude_columns": ["quantity"]
}
}
]
}
left / right: The configuration for each of the two models to compare, along with optional column exclusions.
Dependency Graph (dependency_graph)
- What it checks:
- Monitors the DAG shape for high coupling. It tracks:
fan_in(Inward Coupling): The number of upstream models that this model directly depends on.fan_out(Outward Coupling / Blast Radius): The number of downstream models that depend on this model.
- How to configure:
Defined under
checks.dependency_graphinfitness_functions.yaml. fan_out_warn(int, default: 15): Warn if a model's fan-out is higher than this value.fan_out_fail(int, default: 25): Fail (raise an error) if a model's fan-out is higher than this value.fan_in_warn(int, default: 10): Warn if a model's fan-in is higher than this value.
Materialization Depth (materialization_depth)
- What it checks:
- Calculates the nesting depth of SQL models materialized as
view. Views built on other views incur overhead. - A view's depth is calculated recursively:
1 + max(depth of view dependencies). - Non-views (e.g.,
table,incremental, orseedmodels) reset the depth calculation and have a depth of0. - How to configure:
Defined under
checks.materialization_depthinfitness_functions.yaml. max_depth_warn(int, default: 3): Warn if nesting depth exceeds this.max_depth_fail(int, default: 5): Raise an error if nesting depth exceeds this.
Duplicate CTEs (duplicate_ctes)
- What it checks:
- Identifies "Connascence of Algorithm" by flagging duplicate or near-identical transformation logic inside CTEs across different models.
- CTEs are parsed, canonicalized using
sqlglotto ignore whitespace/formatting differences, and hashed. - CTE extraction and fingerprinting run concurrently across a parallel worker pool for high-performance traversal on large DAGs.
- Only "complex" CTEs are checked. A CTE is complex if it has a minimum AST node count and contains a structural element (
JOIN,WHERE,GROUP BY,HAVING,WINDOW,CASE, orIF). - Centrally managed CTE transformation logic generated by shared macros (dbt Jinja macros, SQLMesh macros, or Dataform functions) is automatically detected and excluded from duplication findings when
ignore_macros: true(default). - How to configure:
Defined under
checks.duplicate_ctesinfitness_functions.yaml.
Connascence of Value (connascence_of_value)
- What it checks:
- Identifies "Connascence of Value" by flagging literal values (strings, numbers) duplicated across multiple models.
- Only domain-meaning literals are checked. Structural and technical SQL literals are automatically excluded based on AST context:
- Literals inside
LIMITorOFFSETclauses. - Data type parameters and precision/scale definitions (e.g.
DECIMAL(15, 2),VARCHAR(255)). - Rounding and truncation precision/scale arguments (e.g.
ROUND(amount, 2),TRUNC(amount, 2)). - String splitting delimiter and index arguments (e.g.
SPLIT_PART(email, '@', 2)). - Positional string slicing parameters (e.g.
SUBSTRING(name, 1, 10),LEFT(name, 5),RIGHT(name, 5)). - String concatenation operators (
||/DPipe) andCONCAT_WSseparators. - Mathematical divisors in arithmetic division (e.g.
amount / 100.00cents-to-dollars divisor).
- Literals inside
- Short punctuation characters (
|,,-,_,/,:) are ignored by default and configurable viaignored_punctuation. - Tiny string literals can be excluded by setting
min_length(e.g.min_length: 3to ignore'0','-1','','3'). - Project-specific literal escapes can be added to
ignored_values. - Grouping is case-insensitive for strings, but the original casing is preserved in the findings messages.
-
How to configure: Defined under
checks.connascence_of_valueinfitness_functions.yaml.checks: connascence_of_value: enabled: true severity: warning # Severity of finding: 'warning' or 'error' min_occurrences: 2 # Minimum number of unique models sharing a literal to trigger (default: 2) min_length: 0 # Minimum length for string literals to check (default: 0) ignored_values: ["0", "1", ""] # List of literals to ignore (default: ['0', '1', '']) ignored_punctuation: ["|", " ", "-", "_", "/", ":"] # Punctuation strings to ignore (default: ['|', ' ', '-', '_', '/', ':']) skip_layers: [staging] -
Why it matters (The "Why"): Connascence of Value occurs when two or more components must share a specific value (literal/constant) to function correctly. If that value changes in the source data or business rules (e.g.
'premium_tier'becomes'premium_membership'), all models containing it must be updated simultaneously. If any are missed, it silently introduces data discrepancies between your models (e.g. marketing counts new users but finance continues to filter on the old tier name). -
Example of Duplication:
-
How to Resolve:
- Upstream Classification (Recommended): Evaluate and rename/classify the status once in a staging layer, exposing it downstream as a simple boolean flag:
- Classification Macros: Encapsulate the comparison predicate or status value into a reusable macro, transforming Connascence of Value into weaker Connascence of Name:
- Project-Level Variables: Define the value as a project variable (e.g. in
dbt_project.ymlor SQLMesh config) for environment configurations or global thresholds, and reference it via Jinja: - Mapping Tables (Seeds): For larger sets of constants or multi-attribute categories (e.g., list of VIP email domains, country code lookups), load them via a seed CSV and perform a
JOINorWHERE IN (SELECT ... FROM {{ ref('seed') }}).
Join Type Parity (join_type_parity)
- What it checks:
- Validates data type parity for joined columns across SQL queries to eliminate Connascence of Type (CoT).
- Walks join condition expressions (
ON left.col = right.colandUSING (col)), extracts columns on both sides, resolves their data types from project model metadata, and flags any incompatible type mismatches. -
Honors explicit casts (e.g.
CAST(id AS VARCHAR) = user_id) and recognizes dialect-equivalent type families (e.g.VARCHAR$\leftrightarrow$TEXT,INT$\leftrightarrow$BIGINT). -
Why it matters (The "Why"): Joining columns with mismatching data types (e.g.
VARCHARjoined toINTEGER) causes dynamic type casting overhead across every processed row, disables database index and partition scans, or leads to runtime execution failures. It introduces Connascence of Type (CoT) where models are tightly and invisibly coupled to upstream internal type representations. -
How to configure: Defined under
checks.join_type_parityinfitness_functions.yaml:checks: join_type_parity: enabled: true severity: error # 'error' or 'warning' skip_layers: [staging] # Optional: layers to skip equivalent_types: # Optional: custom equivalent type families text: [text, varchar, string, char, nvarchar] integer: [int, integer, bigint, smallint, tinyint] numeric: [decimal, numeric, number] float: [float, double, real] timestamp: [timestamp, timestamptz, timestamp_ntz, datetime] -
Example Violation:
-
How to Resolve:
- Align Upstream Schema (Recommended): Cast or define the column with the correct canonical type in the upstream staging model.
- Explicit Casting: If disparate types are intentional, add an explicit
CASTat the join condition:
2. Linter Rules
Linter rules inspect individual model files to enforce code style, conventions, and database-independent references.
For SQLMesh projects, these rules run dynamically inside SQLMesh (e.g., sqlmesh lint) using the lowercase class name.
Ban SELECT * (ban_select_star)
- What it checks:
- Disallows the use of wildcard
SELECT *statements. Requires explicit column naming to reduce model coupling. Aggregate count expressions (e.g.,COUNT(*),COUNT(DISTINCT *)) are permitted. - How to configure:
Defined under
rules.ban_select_starinfitness_functions.yaml. - SQLMesh Rule Name:
banselectstar - Default
skip_layers:["sources"] - Auto-fix: Supported via
tff lint --fixwhen upstream relation schema or CTE projection is statically available in the project context. Uncataloged sources are gracefully skipped.
No Positional GROUP BY/ORDER BY (no_positional_group_by_or_order_by) [Auto-fixable]
- What it checks:
- Prevents using ordinal integers (e.g.,
GROUP BY 1, 2orORDER BY 1 DESC) instead of explicit column name references. - Can be configured together under
rules.no_positional_group_by_or_order_by, or individually viarules.no_positional_group_byandrules.no_positional_order_by. - How to configure:
Defined under
rules.no_positional_group_by_or_order_byor individual rule sections infitness_functions.yaml.rules: # Option 1: Configure together with individual toggles no_positional_group_by_or_order_by: enabled: true group_by: true order_by: true skip_layers: [sources] # Option 2: Configure separately no_positional_group_by: enabled: true skip_layers: [sources] no_positional_order_by: enabled: true skip_layers: [sources] - Checks:
no_positional_group_by(nopositionalgroupby): flags positionalGROUP BYreferences.no_positional_order_by(nopositionalorderby): flags positionalORDER BYreferences.no_positional_group_by_or_order_by(nopositionalgroupbyororderby): legacy umbrella check name.
- Default
skip_layers:["sources"]
Environment Agnostic References (environment_agnostic_references)
- What it checks:
- Blocks hardcoded references to specific database catalog or schema names (like
prod.database.table). - Table references are parsed, and the non-table prefixes are scanned case-insensitively against the banned list.
- How to configure:
Defined under
rules.environment_agnostic_referencesinfitness_functions.yaml. - SQLMesh Rule Name:
environmentagnosticreferences banned_environments(list of strings, default:["prod", "dev", "staging", "uat", "qa"]): Environment strings to block.
Classification Macros (classification_macros)
- What it checks:
- Enforces "Connascence of Meaning" by requiring classification columns to use standard macros instead of inline
CASEstatements. - If a query defines an inline
CASE ... END AS <column>matching a key incolumns, it flags a violation unless the corresponding macro pattern is matched in the query. - How to configure:
Defined under
rules.classification_macrosinfitness_functions.yaml. - SQLMesh Rule Name:
classificationmacros columns: A dictionary mapping target column names to regex patterns matching their expected macro representations.- Default
skip_layers:["sources"]
SQL Complexity (sql_complexity)
- What it checks:
- Evaluates maintainability metrics of a model query:
cte_count: Number of common table expressions.join_count: Number ofJOINstatements.line_count: Total lines of code (ignoring empty lines and SQLMeshMODELblocks).decision_points: Number of logical conditional statements (CASE,IFand boolean operatorsAND/ORinWHEREclauses).nested_subquery_in_final_select: Warns if a subquery is nested in the final SELECT statement FROM or JOIN clauses. (Auto-fixable viatff lint --fix)
- How to configure:
Defined under
rules.sql_complexityinfitness_functions.yaml. - SQLMesh Rule Name:
sqlcomplexity thresholds: Map of metric to[warn_threshold, fail_threshold]integer pairs. Exceedingwarn_thresholdraises a warning; exceedingfail_thresholdraises an error.severity(string, optional, default:"error"): Overall rule severity override ("warning"or"error"). Set to"warning"if you want all complexity breaches to be emitted as warnings only and never fail the run.- Note: The legacy
warn_onlysetting is deprecated and no longer supported; useseverity: warningor threshold pairs instead.
Mart Naming (mart_naming)
- What it checks:
- Enforces naming conventions for models residing inside subfolders of the
martslayer directory. - Ensures that the filename starts with the name of the subfolder directory (e.g.,
marts/marketing/ad_performance.sqlshould be namedmarketing_ad_performance.sql). - How to configure:
Defined under
rules.mart_naminginfitness_functions.yaml. - SQLMesh Rule Name:
martmodelnamingconvention layer_name(string, default:"marts"): Folder name of the marts layer.rule(string, default:"prefix_with_subdirectory"): Naming rule to enforce.
Column Names (column_names)
- What it checks:
- Enforces naming standards on column columns by checking for deprecations or forbidden substrings.
- How to configure:
Defined under
rules.column_namesinfitness_functions.yaml. - SQLMesh Rule Name:
columnnames replacements: A dictionary mapping search regex patterns (deprecated names) to target replacement suggestions (applied viare.sub).
Column Types (column_types)
- What it checks:
- Ensures columns matching specific name patterns are defined with expected data types (e.g., columns ending in
_idmust be typed astext). - How to configure:
Defined under
rules.column_typesinfitness_functions.yaml. - SQLMesh Rule Name:
columntypes rules: A list of rule entries containing:name: Identifier of the rule.pattern: Regex matching column names.data_type: Expected SQL data type.
equivalent_types: A dictionary of synonym types mapping an expected type to list of accepted equivalent strings.
Metadata (metadata) [Partially Auto-fixable]
- What it checks:
- Enforces model metadata documentation and testing:
owner[Auto-fixable]: Validates that the model config has a specified owner.description[Auto-fixable]: Validates that the model description is defined and non-empty.grain: Validates that grains (primary key/grain definition) are specified.not_null: Validates that the model has anot_nullaudit (SQLMesh) or test (dbt).unique_values: Validates that the model has aunique_valuesaudit (SQLMesh) oruniquetest (dbt).
- How to configure:
Defined under
rules.metadatainfitness_functions.yaml. - SQLMesh Rule Names: Runs as five separate rules:
nomissingownernomissingdescriptionnomissinggrainnomissingnotnullnomissinguniquevalues
Filename Equals Model Name (filename_equals_modelname)
- What it checks:
- Validates that the model's catalog identifier matches the stem of its source SQL file on disk.
- How to configure:
Defined under
rules.filename_equals_modelnameinfitness_functions.yaml. - SQLMesh Rule Name:
filenameequalsmodelname
3. Health Scoring Configuration
tff calculates an overall architecture health score (0–100) aggregated from all executed checks. By default, every check carries equal weight (1.0), and failures subtract penalties proportionally (an error penalty of 1.0 and warning penalty of 0.5 per affected model; or 100.0 error and 50.0 warning for project-level checks).
You can configure custom weights and failure penalties under the health: section in fitness_functions.yaml.
Check and Category Weights
Assign custom relative weights to prioritize specific quality dimensions. Checks with higher weights have a greater influence on the overall score.
health:
weights:
layer_integrity: 3.0 # Higher weight for critical architectural boundaries
schema_contracts: 2.0
column_names: 0.5 # Lower weight for naming conventions
metadata: 1.5
# Optionally set weights by connascence category
category_weights:
dynamic_coupling: 2.0 # Connascence of Timing / Execution
static_coupling: 1.5 # Connascence of Position / Meaning
naming: 0.75 # Connascence of Name
- Check-level weights (
weights): Map check or rule names (e.g.layer_integrity,ban_select_star) to a positive float weight. - Category weights (
category_weights): Map categories (e.g.connascence_of_algorithm,dynamic_coupling,metadata) to a positive float weight. If both category and check weights are specified, check-level weights take precedence.
Failure Penalties
Customize the penalty points deducted for errors and warnings:
health:
penalties:
# Model-level penalties (deducted proportionally to model count)
error: 1.0 # Default: 1.0
warning: 0.5 # Default: 0.5
# Project-level penalties (subtracted directly from the check's 100-point score)
project_error: 100.0 # Default: 100.0 (or decimal 1.0)
project_warning: 50.0 # Default: 50.0 (or decimal 0.5)
# Check-specific overrides
checks:
schema_contracts:
error: 2.0 # Strict penalty for schema mismatch
column_names:
warning: 0.1 # Mild penalty for column name warnings
Scoring Formula
-
Model-Level Checks (e.g.
ban_select_star,metadata): $$\text{penalty points} = (\text{error count} \times \text{penalty}{\text{error}}) + (\text{warning count} \times \text{penalty})$$ $$\text{score} = \max\left(0, 100 \times \left(1 - \frac{\text{penalty points}}{\text{penalty}_{\text{error}} \times \text{total models}}\right)\right)$$} -
Project-Level Checks (e.g.
layer_integrity,dependency_graph): $$\text{score} = \max\left(0, 100 - (\text{error count} \times \text{penalty}{\text{proj_error}} + \text{warning count} \times \text{penalty})\right)$$} -
Overall Health Score: The weighted average across all active checks: $$\text{Overall Score} = \frac{\sum (\text{score}_i \times \text{weight}_i)}{\sum \text{weight}_i}$$
4. Custom Plugins & Extensions
tff provides an extensible plugin architecture that enables teams to implement proprietary fitness rules, custom architectural DAG checks, and third-party pipeline adapters without modifying tff-core.
[!TIP] For a complete, step-by-step authoring guide with AST traversal, DAG collector functions, and custom adapter scaffolding, see the dedicated Extending tff Guide.
Overview
Plugins can be registered through two mechanisms:
1. Python Package Entry Points: Distributed Python packages exposing tff.rules and tff.adapters entry points are discovered automatically when installed in your Python environment.
2. Configuration File Plugins: Local Python scripts or installed modules declared under the plugins: list in fitness_functions.yaml.
# fitness_functions.yaml
plugins:
- rules/my_custom_rules.py
- my_company_tff_plugin
rules:
company_naming_convention:
enabled: true
severity: error
prefix: "corp_"
Authoring Custom Rules
To define a custom model-level rule, inherit from tff.core.rules.base.Rule and implement check_model:
# rules/my_custom_rules.py
from tff.core.model import ModelRepresentation
from tff.core.rules.base import Rule, RuleViolation
class CompanyNamingRule(Rule):
"""Enforce that production models follow internal company naming standards."""
name = "company_naming_convention"
category = "Internal Standards"
default_severity = "error"
def check_model(self, model: ModelRepresentation) -> RuleViolation | None:
cfg = self.get_rule_config() or {}
prefix = cfg.get("prefix", "corp_") if isinstance(cfg, dict) else getattr(cfg, "prefix", "corp_")
if not model.name.startswith(prefix):
return self.violation(f"Model '{model.name}' must start with required prefix '{prefix}'.")
return None
tff automatically discovers and registers any Rule subclasses found within files or modules listed in plugins:. Custom configuration options can be retrieved dynamically via self.get_rule_config().
Module Registration Hooks (Optional)
If you prefer explicit control, you can define a register or register_rules hook function in your plugin module:
from tff.core.registry import CheckDefinition, CheckRegistry
def register(registry: CheckRegistry) -> list[CheckDefinition]:
return [
CheckDefinition(
id="custom_dag_check",
label="Custom DAG Validation",
category="Custom",
scope="dag",
collector_module="my_plugin.collector",
collector_func_name="run_dag_check",
)
]
Authoring Custom Pipeline Adapters
Third-party adapters allow tff to analyze non-standard data pipelines or internal orchestration frameworks. Subclass tff.core.adapter.PipelineAdapter:
from pathlib import Path
from tff.core.adapter import PipelineAdapter, register_adapter
from tff.core.model import ModelRepresentation
class MyCustomEngineAdapter(PipelineAdapter):
@property
def provider_name(self) -> str:
return "custom_engine"
def is_applicable(self, project_root: Path) -> bool:
"""Return True if project_root contains this engine's project definition."""
return (project_root / "custom_pipeline.yml").is_file()
def load_models(self, project_root: Path, dialect=None, manifest_path=None) -> dict[str, ModelRepresentation]:
# Parse and return ModelRepresentation mapping
return {}
def run_checks(self, project_root: Path, config, checks=None, dialect=None, manifest_path=None, models=None):
# Run checks or delegate to CheckRegistry
return [], 0, []
To register the adapter, either:
- Expose it in a plugin file/module (auto-discovered),
- Define register_adapters() in your plugin module, or
- Call register_adapter("custom_engine", MyCustomEngineAdapter).
Registering via Python Package Entry Points
For installable libraries (e.g. distributed via PyPI or internal wheels), configure entry points in pyproject.toml:
[project.entry-points."tff.rules"]
company_naming = "my_package.rules:CompanyNamingRule"
custom_checks = "my_package.rules:register"
[project.entry-points."tff.adapters"]
custom_engine = "my_package.adapter:MyCustomEngineAdapter"
Once installed, tff discovers and registers these rules and adapters automatically without requiring explicit plugins: configuration in fitness_functions.yaml.
Inspecting Plugins and Adapters
You can inspect registered plugins and adapters using the tff info command: