Data Scientist (2026-2201)

McLean, VA – Full Time

RESPONSIBILITIES

  • Engage with customers to determine the nature of requirement/analytic problem, evaluate options, and offer Information Technology (IT)-based recommendations/options

  • Advise customers on IT methods and data needed and/or available to satisfy the requirement

  • Develop IT solutions to enable the exploitation of databases and files; devise creative data collection and strategies

  • Develop comprehensive software applications

  • Identify potential additional data sources and automate collection processes

  • Integrate, analyze, evaluate, and assess quantitative data (using statistical software, computer models, geospatial models, software languages, mathematical models, machine learning) to contribute to or develop IT solutions to requirements (software tools, analytic models, or reports.)

  • Conduct statistical, mathematical, geospatial modeling or data-mining analysis in partnership with other colleagues

  • Anticipate and project a wide range of possible outcomes using scenario/alternative analyses, machine learning, and agent based modeling or other advanced analytic techniques

  • Identify, use, and/or develop a wide range of methodologies and analytic tools to address existing or potential problems and collection strategies

  • Prepare, and communicate a wide range of strategic, highly complex graphics, computational models/tools, or written/oral assessments

  • Leverage multiple data management tools to organize relevant information and make decisions

REQUIRED SKILLS, EXPERIENCE & QUALIFICATIONS

  • Active TS/SCI clearance with Poly.

  • Demonstrated experience building production data pipelines and ETL/ELT workflows at scale. 

  • Demonstrated experience with Apache Spark and PySpark for distributed data processing. 

  • Demonstrated experience with advanced Python programming skills including data manipulation libraries (Pandas, NumPy) and data engineering best practices. 

  • Demonstrated experience understanding data security, privacy, governance, and compliance principles. 

  • Demonstrated experience with workflow orchestration tools such as Step Functions and Airflow. 

  • Demonstrated experience with containerization such as Docker or Podman, and deploying data applications in cloud environments. 

  • Demonstrated experience with AWS services (in particular S3, Lambda, and Step Functions). 

  • Demonstrated experience with PostgreSQL and MySQL in production environments, including performance tuning and schema design. 

  • Demonstrated experience with SQL and query optimization for complex analytical workloads. 

  • Demonstrated experience with version control (Git) and CI/CD practices for data pipelines. 

  • Demonstrated experience working with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight. 

  • Demonstrated experience with strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks. 

PREFERRED SKILLS, EXPERIENCE & QUALIFICATIONS

  • Demonstrated experience with data lakehouse architectures using Apache Iceberg.

  • Demonstrated experience configuring, deploying, and integrating data platform components.

  • Apache Ranger (access control and data governance)

  • Trino (distributed SQL query engine)

  • Apache Superset (data visualization and dashboarding)

  • Demonstrated experience with Bash scripting for automation and data processing tasks.

  • Demonstrated experience with Infrastructure as Code (Terraform or CloudFormation) for data infrastructure.

  • Demonstrated experience with tracking data lineage and associated tooling such as OpenLineage.

  • Demonstrated experience with Java.

  • Demonstrated experience with data quality frameworks, testing methodologies, and validation strategies.

  • Demonstrated experience or background with large-scale data migrations or platform modernization efforts.

  • Demonstrated experience integrating AI/ML services and models (translation, OCR, speech-to-text, NLP, language detection, topic modeling), LLMs, and RAG (retrieval-augmented generation) pipelines.

  • Demonstrated experience with geospatial data processing (H3, PostGIS, or similar).

  • Demonstrated experience Contributing to data engineering documentation, best practices, or design patterns.

  • Demonstrated experience with NoSQL databases (DynamoDB, etc.).

  • Demonstrated experience with excellent written and verbal communication skills with both technical and non-technical audiences.