View all jobs
Senior Data Engineer — Azure / Fraud Analytics
Location: Remote / Telework (U.S. based) · Occasional on-site collaboration as requested Job Type: Full-time Clearance: Must be able to obtain and maintain a Public Trust background investigation Citizenship: U.S. Citizenship required
You’ll work code-first — source control, CI/CD, and reusable modular design are core to how this team operates. If you enjoy owning architecture end to end and partnering with data scientists to make ML pipelines fast and reliable, this role is for you.
About the Role
We are hiring a Senior Data Engineer to build and maintain the Azure-based data architecture that powers fraud analytics for a federal Office of Inspector General (OIG). You will design integrated, flexible, and secure pipelines that move source data into a modern cloud environment and make it available for audits, investigations, and machine learning.You’ll work code-first — source control, CI/CD, and reusable modular design are core to how this team operates. If you enjoy owning architecture end to end and partnering with data scientists to make ML pipelines fast and reliable, this role is for you.
What You’ll Do
- Design, implement, and maintain an efficient, secure, stable, and flexible data architecture, with all assets managed via source control
- Migrate source data to Azure Data Lake Storage (ADLS) and build/maintain ELT/ETL pipelines in Azure Synapse and Azure Machine Learning (SDK V1 and V2)
- Review and improve existing architecture and pipelines — periodic audits to address bottlenecks, deprecated dependencies, and architecture drift
- Establish quality controls, error handling, logging, and validation checks across all pipelines
- Normalize entity attributes (addresses, phone numbers, and other common fields)
- Optimize ingestion, processing, and storage across diverse datasets, including modern columnar formats such as Parquet
- Build self-service capabilities so analysts can query and export data for investigations and audits
- Partner with data scientists to ensure the architecture efficiently supports machine learning workloads
- Author and maintain SOPs governing authoring, validation, publishing, execution, and monitoring of all pipelines and assets
- Produce detailed documentation — data dictionaries, ER diagrams, and pipeline process maps
- Stay current with emerging AI tooling and contribute to efforts evaluating automation and LLM-assisted capabilities
Required Qualifications
- 5 years of hands-on experience in each of:
- Maintaining SQL databases and conducting advanced operations in SQL and T-SQL
- Designing, implementing, and maintaining ELT/ETL processes in cloud-based data analytics environments
- 3 years of hands-on experience in each of:
- Working in Azure Synapse and Azure Machine Learning with the modern data stack — certifications preferred (DP-203 or equivalent)
- Manipulating data in Python (Pandas required; PySpark/Polars preferred; experience with reusable, modular code preferred)
Preferred
- Implementing pipelines and infrastructure using code-first approaches (Python SDK, CLI, REST APIs, or IaC tooling)
- Implementing source control and CI/CD workflows
- Demonstrated familiarity with AI coding assistants and LLM integration patterns
- Azure certification (DP-203 or equivalent)
Work Schedule & Environment
- Core hours between 6:00 a.m. and 6:00 p.m. local time, Monday–Friday
- Telework authorized; occasional on-site meetings and collaboration as requested
- Government-furnished equipment provided for work on government systems
- Federal holidays observed
