$142,320 – $213,480
Listed on Citigroup’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This VP-level data engineering leadership role at Citi involves architecting enterprise data platforms and leading a team to build production-grade pipelines supporting AI and analytics at scale. It suits experienced data engineers ready to move into strategic technical leadership, designing infrastructure for machine learning and real-time processing in a major financial services environment.
Our summary, not Citigroup’s wording. The full posting is on their site.
Skills this role names
- Apache Airflow
- Apache Flink
- Apache Hive
- Apache Iceberg
- Apache Kafka
- Chroma
- Databricks
- Docker
- Java
- Kubernetes
- Milvus
- Pinecone
- Python
- Qdrant
- SQL
- Terraform
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's degree in Computer Science, Data Engineering, Information Systems, or quantitative field
- 6+ years professional experience in data engineering, software engineering, or data platform development
- 3+ years in technical lead or engineering leadership role
- Expert-level Python and SQL proficiency
- Deep expertise in Apache Iceberg, Apache Pinot, Hive/HDFS, and distributed computing
- Hands-on experience building real-time streaming pipelines with Apache Kafka and Apache Flink at scale
- Experience designing data infrastructure for machine learning, analytics, or AI in production
- Experience with modern table formats like Delta Lake, Apache Iceberg, or Apache Hudi
- Experience with workflow orchestration tools such as Apache Airflow
Nice to have
- Master's degree in Computer Science, Data Engineering, Information Systems, or related quantitative discipline
- Additional Java proficiency
- Experience with Databricks including cluster tuning and lakehouse optimization
- Familiarity with vector databases such as Milvus, Pinecone, Qdrant, or Chroma for GenAI and LLM use cases
- Understanding of dimensional modeling, Data Vault, and schema-on-read/write design patterns
- Hands-on experience with Docker, Kubernetes, and Terraform for Infrastructure as Code