Engineering Reliable and Scalable Feature Pipelines for Quantitative Finance

Quantitative investment systems depend on data pipelines whose errors can silently propagate into model features, rankings, and decisions. During a six-month Data Engineering internship at Axyon AI, I reconstructed and validated a commodity-futures pipeline, corrected stateful rolling logic in the production dbt/Snowflake stack, catalogued feature definitions with Feast, and scaled genome-driven feature computation through Polars concurrency, PySpark, and remote AWS execution. The unifying objective was to improve reliability, governability, and scalability while preserving the business meaning of existing features.
The reliability work separated calendar state from price availability in the continuous-futures model. Expiration and rolling events could fall on holidays, with missing prices; filtering those rows before computing the cumulative state resulted in incorrect contract selection for four commodities (HG, SI, LA12, AL). The fix computes state on a stable calendar, conditionally forward-fills event-day prices, and propagates the resulting effective prices into lag and return calculations — recomputing the roll/expiry phase for roughly 22,000 daily commodity-date rows, which this report treats as the headline result rather than an accuracy percentage. A controlled Postgres/Docker/Airflow implementation and a row-level comparison framework against Snowflake support this work; the framework is reused in Chapter 4 to evaluate whether SQLMesh could replace dbt, through key coverage and MAPE.
The governance work extended Feast from a legacy commodity view to Datasmith 2.0 transformations and YAML genomes: transformations became FeatureViews, genome compositions became FeatureServices, and a metadata identity based on transformation group, method, and exact parameters prevented duplicate catalogue entries even when aliases differed. The registry was moved to Snowflake and exposed to Talos for registration and retrieval.
The scalability work introduced two bounded levels of parallelism — independent feature groups, and dependency-aware feature waves within each group — which exposed shared singleton state, resolved through locks and thread-local feature-group type. An optional Spark backend partitioned the target computation by instrument, reused the existing Polars formulas, reproduced the cross-sectional ranking in Spark, and wrote the output without collecting the full frame on the driver; the experiment was then connected to Malphas/AWS Batch for remote execution. Locally, eight-feature workers were slower than two (119.6 s versus 97.9 s), indicating that worker count must be tuned, as Polars already multithreads internally. The principal result is therefore not a single speedup number, but a tested engineering method: isolate state from availability, make definitions explicit, parallelise only independent work, and validate every alternative against a trusted reference.