Paytm
About this role
This is big data engineering on Paytm's insurance vertical, working on the platform that carries
petabytes a day. The job is building and running the pipelines: real time event ingestion for live
dashboards, transformations that turn raw sources into reliable components of the data lake, and new
components on the Hadoop ecosystem. The stack named is Hadoop, MapReduce, Hive, Spark and PySpark,
with Python, Java or Scala for the code, and Kafka, HBase or Cassandra, Redis and AWS S3 around the
edges. Note the posting contradicts its own title: the board calls it "Senior / Lead Engineer" while
the body's first line reads "Data Engineering, Technical Lead". The requirement list is purely
technical with no team management in it. Apply if you want data volume that is genuinely hard.
petabytes a day. The job is building and running the pipelines: real time event ingestion for live
dashboards, transformations that turn raw sources into reliable components of the data lake, and new
components on the Hadoop ecosystem. The stack named is Hadoop, MapReduce, Hive, Spark and PySpark,
with Python, Java or Scala for the code, and Kafka, HBase or Cassandra, Redis and AWS S3 around the
edges. Note the posting contradicts its own title: the board calls it "Senior / Lead Engineer" while
the body's first line reads "Data Engineering, Technical Lead". The requirement list is purely
technical with no team management in it. Apply if you want data volume that is genuinely hard.
Who this is for
Requirements. 4 to 8 years of experience in big data technologies. Strong hands-on Hadoop,
MapReduce, Hive, Spark and PySpark. Excellent programming and debugging in Python, Java or Scala.
Experience with a scripting language such as Python or Bash. Hands-on programming with multithreaded
applications. The posting also asks for a strong product design sense and specialisation in Hadoop
and Spark specifically.
Good to have. NoSQL databases such as HBase or Cassandra. Databases, SQL and messaging queues like
Kafka. Streaming applications using Spark Streaming, Flink or Storm. AWS and cloud technologies such
as S3. Caching architectures such as Redis.
The work itself. Grow the analytics capability with faster and more reliable tooling, handling
petabytes of data every day. Build platforms that make data available to cluster users with low
latency and horizontal scalability. Diagnose problems across the whole technical stack. Design and
develop a real time events pipeline for data ingestion feeding real time dashboards. Write complex
transformation functions turning raw sources into reliable data lake components. Design and implement
new components across the Hadoop ecosystem.
Title contradiction, worth knowing before you apply. The job board title is "Data Engineering,
Senior / Lead Engineer" and the first line of the description is "Data Engineering, Technical Lead".
Nothing in the responsibilities or requirements describes managing people, so we have published it at
the stated 4 to 8 year band as an individual contributor role. If the level matters to you, ask in
the first call which of the two the requisition actually is.
Location and terms. Noida, marked on-site, Full-time Employment, on the Insurance team.
Who this is for. A data engineer with four or more years on the Hadoop and Spark side who wants
volume. If your background is warehouse and SQL rather than distributed processing, the mandatory
list here is squarely the latter.
MapReduce, Hive, Spark and PySpark. Excellent programming and debugging in Python, Java or Scala.
Experience with a scripting language such as Python or Bash. Hands-on programming with multithreaded
applications. The posting also asks for a strong product design sense and specialisation in Hadoop
and Spark specifically.
Good to have. NoSQL databases such as HBase or Cassandra. Databases, SQL and messaging queues like
Kafka. Streaming applications using Spark Streaming, Flink or Storm. AWS and cloud technologies such
as S3. Caching architectures such as Redis.
The work itself. Grow the analytics capability with faster and more reliable tooling, handling
petabytes of data every day. Build platforms that make data available to cluster users with low
latency and horizontal scalability. Diagnose problems across the whole technical stack. Design and
develop a real time events pipeline for data ingestion feeding real time dashboards. Write complex
transformation functions turning raw sources into reliable data lake components. Design and implement
new components across the Hadoop ecosystem.
Title contradiction, worth knowing before you apply. The job board title is "Data Engineering,
Senior / Lead Engineer" and the first line of the description is "Data Engineering, Technical Lead".
Nothing in the responsibilities or requirements describes managing people, so we have published it at
the stated 4 to 8 year band as an individual contributor role. If the level matters to you, ask in
the first call which of the two the requisition actually is.
Location and terms. Noida, marked on-site, Full-time Employment, on the Insurance team.
Who this is for. A data engineer with four or more years on the Hadoop and Spark side who wants
volume. If your background is warehouse and SQL rather than distributed processing, the mandatory
list here is squarely the latter.
Apply on company site
Opens jobs.lever.co, the employer's own application page. Applying is always free.