Infinite
Texas

Data Engineer

OnsitePosted yesterday

We tailor your resume to this role and apply for you in seconds.

Apply to Data Engineer at Infinite

Job details

Location
Texas
Work type
Onsite
Posted
yesterday
Apply on
sjobs.brassring.com

About this role

Job description

Role Summary

We are seeking an experienced Big Data Engineer to design, build, and optimize large-scale batch and streaming data pipelines on the Hadoop ecosystem using Apache Spark and Scala. The role supports high-volume ingestion, transformation, and enrichment of clickstream, network, and location datasets, working closely with data architects, platform engineering, and downstream analytics teams. This is a hands-on engineering role with ownership of pipeline performance, reliability, and data quality in production.

Key Responsibilities

  • Design, develop, and maintain distributed data pipelines using Apache Spark (Core, SQL, Streaming) written in Scala.
  • Build ingestion and transformation workflows across the Hadoop ecosystem — HDFS, Hive, YARN, MapReduce — for structured and semi-structured data at TB–PB scale.
  • Tune and optimize Spark jobs: partitioning strategy, caching, broadcast joins, shuffle reduction, data skew handling, and executor/memory sizing.
  • Implement real-time and near-real-time ingestion using Apache NiFi and/or Kafka.
  • Embed data quality, reconciliation, and validation controls directly into pipelines.
  • Author and optimize HiveQL and Spark SQL for curated and consumption layers.
  • Automate orchestration and scheduling using Airflow, Oozie, or Control-M.
  • Participate in code reviews, CI/CD automation, unit and integration testing, and production support.
  • Troubleshoot job failures, SLA breaches, and performance regressions; drive root-cause analysis to permanent fixes.
  • Document data flows, lineage, transformation logic, and operational runbooks.

Required Qualifications

  • 6+ years of data engineering experience, with 4+ years hands-on Apache Spark development in Scala on production workloads.
  • Strong Scala fundamentals — functional programming constructs, collections API, case classes, pattern matching, implicits, and error handling.
  • Deep working knowledge of the Hadoop ecosystem: HDFS, Hive, YARN, HBase.
  • Advanced SQL and data modeling skills across dimensional and big-data denormalized patterns.
  • Demonstrated Spark performance tuning and debugging using the Spark UI, event logs, and physical execution plans.
  • Proficiency with columnar and serialization formats — Parquet, ORC, Avro — including compression and partitioning trade-offs.
  • Linux and shell scripting, Git, Maven or SBT, and Jenkins or equivalent CI/CD tooling.
  • Ability to work independently in a distributed onshore–offshore delivery model.

Preferred Qualifications

  • Kafka and Spark Structured Streaming for event-driven pipelines.
  • Cloud data platform exposure — GCP (BigQuery, Dataproc), AWS EMR, or Azure Databricks.
  • Telecom domain experience with clickstream, network, or geospatial/location data.
  • Python or PySpark as a secondary development language.
  • Data governance and security frameworks — Apache Ranger, Kerberos, PII masking and tokenization

Nice to Have:

  • Apache NiFi flow design, configuration, and administration.
 

Education

Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline — or equivalent demonstrable practical experience.
 

Range of Year Experience-Min Year

15

Physical Location

Texas

Qualifications

Bachelor

Range of Year Experience-Max Year

18
Ready to apply to Infinite?
We tailor your resume to this role and apply for you.

About Infinite

Infinite
Texas