Location : Remote (EST)
Position type : Contract
Job description:
Owns the data pipelines that move data from operational source systems onto the data platform extraction, ingestion, transformation, orchestration, and the day-to-day operational health of those pipelines.
Also plays a role as a hands-on builder of Foundational Data Products (raw record-of-truth) and potentially Derivative Data Products (composed from upstream products).
Â
Core skills
SQL & Python (table stakes); Scala/Java for high-throughput streaming
Pipeline build & operations:Â design, develop, deploy, monitor, and remediate batch and streaming pipelines that land source data on the platform; manage backfills, replays, late-arriving data, schema drift, and SLA breaches
Ingestion patterns:Â full-load, incremental, change-data-capture (Debezium, Fivetran, Qlik Replicate, GoldenGate), event-driven ingest, API and file-based
intake
Pipeline frameworks:Â dbt, Apache Spark, Apache Beam, Airflow, Dagster, Prefect
Streaming:Â Kafka, Kinesis, Flink, Spark Structured Streaming
Data product packaging:Â schema contracts (Avro/Protobuf/JSON Schema), versioning, SLAS/SLOs, data contracts, output-port design (SQL, file, API, event)
Storage formats:Â loeberg, Delta Lake, Hudi (open table formats are now the mesh default)
Quality & observability:Â Great Expectations, Soda, Monte Carlo, dbt tests, data lineage (OpenLineage)
CI/CD for data:Â GitOps pipelines, unit + integration tests, environment promotion
Al-adjacent vector store ingestion (pgvector, Pinecone), feature stores (Feast, Tecton), RAG-ready chunking and embedding pipelines
Mesh-specific:Â knowing when to build a derivative product vs. extending an existing one; consuming upstream products through governed input ports rather than reaching into source systems/
Â
Regards,
Priyanka
Lead Recruiter
Net2Source Inc. |Â Address:Â 270 Davidson Ave, Suite 704, Somerset, NJ 08873, USA
Direct: 201-221 8131 | Email: Priyanka.j@net2source.com Â