1. What is a Data Engineer at Citi?
As a Data Engineer at Citi, you operate at the core of a massive global financial ecosystem. This role is responsible for designing, building, and optimizing the large-scale data pipelines, data warehouses, and streaming architectures that power critical financial products, risk management systems, and trading analytics. Your work directly enables real-time data ingestion, robust governance, and high-performance querying across disparate enterprise sources, ensuring that millions of transactions are processed securely and efficiently every day.
The scale and complexity of Citi's data infrastructure make this position both challenging and strategically vital. You will frequently work with cutting-edge distributed data frameworks—such as PySpark, Kafka, Snowflake, Databricks, and Starburst (Trino/PrestoSQL)—to handle petabyte-scale financial datasets. Whether you are building real-time data acquisition layers for fixed income trading analytics or optimizing complex data lake infrastructures, your solutions directly influence business intelligence capabilities, regulatory compliance, and overall institutional decision-making.
Expect to operate in a fast-paced, highly collaborative environment where technical precision meets stringent financial regulation. You will bridge the gap between raw data generation and actionable insights, partnering closely with software engineers, data scientists, and business stakeholders. Success in this role requires not only deep technical proficiency in distributed data processing and advanced SQL, but also a disciplined approach to data quality, governance, and system performance.

