← Back to jobs

Director, Software Engineer Data Processing Engine Orchestration

Skills

aiai enablementanalyticsawsazurecloudcloud nativedata engineeringdistributed systemsflinkgcpgenerative ailakehousemachine learningspark

Description

Job Summary

DataPelago is at the forefront of revolutionizing data processing for traditional analytics and cutting-edge GenAI preprocessing. We are building an innovative data processing engine that is transforming how Apache Spark, Apache Flink, Ray and others operate on diverse, large-scale data. Our team of engineers drive and adopt advances in hardware-accelerated computing, parallel processing of large-scale data, query optimization, distributed systems, compilers, machine learning, and cloud-native computing. We are looking for specialists to join our engineering team and shape the future of accelerated data processing.

DataPelago Nucleus is a universal data processing engine that is designed to accelerate the processing of diverse data – structured through unstructured – with any parallel processing framework – e.g., Apache Spark – on any infrastructure – including vectorized CPU and GPU. DataPelago Accelerator for Spark (DPA-S), our product based on DataPelago Nucleus, is running large-scale production applications of many globally renowned customers.

The Opportunity:

As senior development leader of DataPelago data processing engine orchestration, you will spearhead the development and delivery of innovative capabilities in the planning, optimization, and coordination of complex workloads on large-scale distributed systems. The orchestration layer drives distributed execution plans adapted to the capabilities of the processing elements in the infrastructure, towards minimizing execution duration and cost. You will be responsible for growing and advancing the development team, development methodologies, and product capabilities.

a key individual contributor in adopting and advancing capabilities of open-source software (OSS) platforms such as Apache Gluten, Velox, Apache Spark, and Apache Flink in the context of DataPelago’s data processing engine. You will enhance functional breadth, performance, scale, and reliability of the DataPelago engine through downstream and upstream contributions. You will have the opportunity to engage with community working on these platforms. This is a unique opportunity to make a significant impact on a category-defining product and work with a talented team of engineers.

What You'll Do:

  • Lead and grow a talented team of engineers with expertise in query planning, query optimization, plan execution orchestration for massively parallel data processing engines.
  • Drive development with technical leads to implement and deliver industry-leading capabilities in data processing engine orchestration.
  • Establish metrics including for product quality, performance, scale, and reliability and lead the team in ensuring these metrics are continuously improved and achieved in customer deployments.
  • Continuously improve software development lifecycle practices leveraging the latest AI advances for higher code quality, release velocity, and productivity.
  • Execute efficiently to fulfill product roadmap aligned with organization goals and priorities.
  • Lead by example and foster a high-performance, high-integrity, collaborative work culture.

What You'll Bring:

  • Bachelor’s degree in computer science or a related field. Master’s or PhD preferred.
  • 12+ years of production software development experience with 7+ years of experience managing software development teams.
  • 10+ years of experience in planning, developing, launching, and supporting core components of massively parallel data processing platforms (e.g., query engines, cloud data warehouses, lakehouse engines) for production deployment.
  • Technical and process rigor in meeting and exceeding data accuracy, high performance, reliability, and security requirements of enterprise data processing products.
  • Ability to attract, retain, and grow exceptional talent in data processing engine orchestration.
  • Deep technical experience in one or more of database systems, data warehouses, query optimization, query engines. Experience with one or more of Apache Spark, Apache Flink, Apache Gluten, Velox, Apache DataFusion, Spark RAPIDS preferred.
  • Solid understanding of the architecture, deployment, and operations of data processing platforms in applications such as data engineering, data preparation, and analytics.
  • Excellent judgment on technology, software development, AI adoption, and people management. Strong leadership skills in all these areas.
  • Good skills engaging with customers, product management, sales, and executive management.

Compensation:
The target salary range for this position is 227,800 - 338,800 USD. The salary offered will be determined by the candidate's location, qualifications, experience, and education and may be outside of this range. Final compensation packages are competitive and in line with industry standards, reflecting a variety of factors, and include a comprehensive benefits package. This may cover Health Insurance, Life Insurance, Retirement or Pension Plans, Paid Time Off, various Leave options, Performance-Based Incentives, employee stock purchase plan, and/or restricted stocks (RSU’s), with all offerings subject to regional variations and governed by local laws, regulations, and company policies. Benefits may vary by country and region, and further details will be provided as part of the recruitment process.

At NetApp, we embrace a hybrid working environment designed to strengthen connection, collaboration, and culture for all employees. This means that most roles will have some level of in-office and/or in-person expectations, which will be shared during the recruitment process.

Equal Opportunity Employer:

NetApp is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination based on age, race, color, gender, sexual orientation, gender identity, national origin, religion, disability or genetic information, pregnancy, protected veteran status, and any other protected classification.

Why You'll Thrive at NetApp

At NetApp, you won't wait for the perfect moment—you'll make it. The early planning, the extra thought, the bold idea that turns good into great: That's how our people operate and how we continue to push the boundaries of data infrastructure.

NetApp is the trusted partner for organizations transforming data into opportunity. As the only enterprise-grade storage service natively embedded in Google Cloud, AWS, and Microsoft Azure, we empower customers to run everything from traditional workloads to enterprise AI with unmatched performance, resilience, and security.

Our culture

We celebrate mold breakers, bold thinkers, and problem solvers. We reward initiative, impact, and ownership. We provide flexibility so you can balance professional ambition with your personal life. Here, differences are not just welcomed—they drive everything we do.

If you're ready to innovate, rise to the challenge, and own every moment - make your next move your best one. Apply now.

Submitting an Application

To ensure a streamlined and fair hiring process for all candidates, our team only reviews applications submitted through our company website. This practice allows us to track, assess, and respond to applicants efficiently. Emailing our employees, recruiters, or Human Resources personnel directly will not influence your application.

AI Disclosure

For select roles, some stages of our hiring process may use artificial intelligence tools to help evaluate applications and candidate selection. These tools support—rather than replace—human decision-making.

Get similar jobs in United States by email

We'll email you when new jobs similar to this one appear.

Similar jobs

Explore more Software Developer jobs in United States.

No similar openings right now. Refine your search or create an alert above.