About this role
This is the senior counterpart to Tekion's other data engineering opening, and the scope is platform rather than pipeline. You would help design a multi tenant, cloud scale data platform serving thousands of dealerships while keeping tenant isolation, governance and data quality intact, and help move the company from batch processing towards near real time and real time ingestion. The technology list is current and specific: Apache Spark, Delta Lake, Apache Iceberg or Apache Hudi, Kafka, Flink or Kinesis, Airflow, and AWS services including EMR, S3, Glue and Athena. It also asks for real data modelling depth, including dimensional modelling and Data Vault. Good fit if you have built warehouses or lakehouses and want the architecture side rather than another set of pipelines.
Who this is for
Asks for 6+ years of experience in data engineering. Required: strong expertise in Python, SQL and Apache Spark; experience building scalable batch and real time ETL and ELT pipelines; hands on experience with AWS services including EMR, S3, Glue and Athena; experience with Kafka, Flink or Kinesis for streaming; strong knowledge of dimensional modelling, Data Vault and data warehousing concepts; and experience with Delta Lake, Apache Iceberg or Apache Hudi.
The responsibilities section also lists: strong experience designing and building scalable data platforms, data warehouses and lakehouse architectures; deep expertise in data modelling including dimensional modelling, Data Vault and enterprise data architecture principles; advanced SQL with query optimisation, performance tuning and large scale data processing; hands on experience with distributed processing frameworks such as Apache Spark; experience designing and implementing batch, streaming and change data capture based ingestion pipelines; proficiency in Python or Scala; and experience with workflow orchestration platforms such as Airflow.
The posting frames the work as building the enterprise data foundation behind analytical products, operational insight and AI features, and describes a modernisation from batch oriented processing to near real time and real time ingestion. Multi tenancy, tenant isolation and governance across thousands of dealerships are called out as the hard part.
Location is Bangalore HQ. No office day count and no interview process are published.
Fit note: if you have only built pipelines inside someone else's platform, the multi tenant isolation and modelling requirements are where this interview will go.
The responsibilities section also lists: strong experience designing and building scalable data platforms, data warehouses and lakehouse architectures; deep expertise in data modelling including dimensional modelling, Data Vault and enterprise data architecture principles; advanced SQL with query optimisation, performance tuning and large scale data processing; hands on experience with distributed processing frameworks such as Apache Spark; experience designing and implementing batch, streaming and change data capture based ingestion pipelines; proficiency in Python or Scala; and experience with workflow orchestration platforms such as Airflow.
The posting frames the work as building the enterprise data foundation behind analytical products, operational insight and AI features, and describes a modernisation from batch oriented processing to near real time and real time ingestion. Multi tenancy, tenant isolation and governance across thousands of dealerships are called out as the hard part.
Location is Bangalore HQ. No office day count and no interview process are published.
Fit note: if you have only built pipelines inside someone else's platform, the multi tenant isolation and modelling requirements are where this interview will go.
Apply on company site
Opens jobs.ashbyhq.com, the employer's own application page. Applying is always free.