Getting Started
Installation
The PyPI distribution is named peh-dataguard, while the Python import
name is dataguard.
Install with uv
For a uv-managed project, add DataGuard as a dependency:
To install it into the current environment without changing a project file:
Install with pip
Create and activate a virtual environment, then install the package:
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install peh-dataguard
On Windows PowerShell, activate the environment with
.venv\Scripts\Activate.ps1 instead of source.
Verify the installation
For a uv-managed project, use uv run python ... instead of python ....
DataGuard workflow
| getting_started.py | |
|---|---|
2 - Instantiate a Validator using config_from_mapping() method.
3 - Validate the dataframe df calling validate()
While this example uses empty lists and an empty DataFrame for simplicity, it illustrates the core three-step process: define your validation rules in a configuration, create a validator from that configuration, and apply it to your data.
Validating real-world constraints
Consider an age column configuration that demonstrates DataGuard's data quality enforcement capabilities:
- Type Safety: Enforcing
integerdata type prevents string or float contamination - Null Prevention:
nullable: Falseensures no missing age values slip through - Range Validation: Age bounds
[0-150)catch unrealistic values like negative ages or extreme outliers - Business Logic: Reflects real-world constraints for human age data
Analyzing error report
Each validation instance that catches validation errors creates an ErrorReport that is
collected to the ErrorCollector under the hood.
Another way to access the ErrorCollector is calling the Validator's error_collector property.
Let's deep dive into the second error Is greater than or equal to.
The DFErrorSchema return provides detailed information about a specific validation error:
DFErrorSchema Fields
-
type: The error category (SchemaErrorReason.DATAFRAME_CHECK) indicating this is a column-level validation failure
-
message: Descriptive error text explaining what failed, including the specific check and example failure cases
-
level: Error severity (ErrorLevel.ERROR) - can be error, warning, or info
-
title: Human-readable error name (Is greater than or equal to) matching the validation rule
-
traceback: Stack trace information (None if not available) for debugging
-
column_names: List of affected columns (['age']) - useful for multi-column validations
-
row_ids: Specific row indices that failed validation ([3]) - enables precise error location
-
idx_columns: Index column information ([]) - empty when not using custom indices
In this example, the error shows that row 3 in the 'age' column contains the value -5, which violates the "greater than or equal to 0" constraint. This granular information allows you to pinpoint exactly which data points need fixing and why they failed validation.