← Back to jobs

Principal Software Engineer, Data Processing Optimization

Skills

aianalyticsawsazurecloudcloud nativedata engineeringdistributed systemsflinkgcpgenerative aimachine learningpythonsaassparksqltrinoPythonAWSGCPAzureSparkSQL

Description

Job Summary

DataPelago is at the forefront of revolutionizing data processing for traditional analytics and cutting-edge GenAI preprocessing. We are building an innovative data processing engine that is transforming how Apache Spark, Apache Flink, Ray and others operate on diverse, large-scale data. Our team of engineers drive and adopt advances in hardware-accelerated computing, parallel processing of large-scale data, query optimization, distributed systems, compilers, machine learning, and cloud-native computing. We are looking for specialists to join our engineering team and shape the future of accelerated data processing.

DataPelago Nucleus is a universal data processing engine that is designed to accelerate the processing of diverse data – structured through unstructured – with any parallel processing framework – e.g., Apache Spark – on any infrastructure – including vectorized CPU and GPU. DataPelago Accelerator for Spark (DPA-S), our product based on DataPelago Nucleus, is running large-scale production applications of many globally renowned customers.

The Opportunity:

As a Senior/Principal Data Processing Optimization Engineer, you will be a key individual contributor in advancing planner and optimization capabilities so that DataPelago acceleration delivers the most performance and cost reduction for open-source software SQL and Python processing engines running on CPU-based and GPU-based infrastructure. This is a unique opportunity to make a significant impact on a category-defining product and work with a talented team of engineers.

What You'll Do:

  • Research, design, implement, and deliver logical and physical query plan optimization techniques for massively parallel data processing engines. Develop query planning and query optimization techniques that enable industry-leading performance and performance efficiency on execution engines based on heterogeneous accelerated compute elements.
  • Integrate DataPelago data processing acceleration and plan optimization capabilities with diverse open-source engines including Apache Spark and Apache Flink.
  • Ensure product performance on customers’ real deployment environment and workload.
  • Establish metrics for plan quality, workload performance, reliability and lead the team in ensuring these metrics are continuously improved and achieved in customer deployments.
  • Provide technical leadership through all phases of product development and post-production support. Own development and support of major components and capabilities.
  • Collaborate with other engineering team members and product management in defining product features, architecture, interfaces, and development timelines. Mentor junior team members.

What You'll Bring

  • Bachelor’s degree in computer science. Master’s or PhD preferred, especially in areas related to database systems.
  • 15+ years of experience developing query planning and query optimization techniques for enterprise data processing software platform. Prefer at least 5+ years of that experience on parallel platforms such as Apache Spark, Presto/Trino, or Apache Flink.
  • Deep knowledge of SQL semantics and expertise in SQL processing. Additional Python processing preferred. Solid understanding of databases, data warehouses, query engines.
  • Deep knowledge of query optimization state of the art, from published literature and open-source software. Prefer own open-source contributions and/or publications in this area.
  • Solid understanding of the architecture, deployment, and operations of data processing platforms in applications such as data engineering, data preparation, and analytics.
  • Demonstrated proficiency in all phases of software development and post-production support of enterprise software or SaaS. Adept in adopting AI effectively.
  • Experience owning the definition and delivery of major product capabilities through technical leadership and collaboration spanning multiple teams.
  • Strong collaboration and communication skills.

Compensation:
The target salary range for this position is 227,800 - 338,800 USD. The salary offered will be determined by the candidate's location, qualifications, experience, and education and may be outside of this range. Final compensation packages are competitive and in line with industry standards, reflecting a variety of factors, and include a comprehensive benefits package. This may cover Health Insurance, Life Insurance, Retirement or Pension Plans, Paid Time Off, various Leave options, Performance-Based Incentives, employee stock purchase plan, and/or restricted stocks (RSU’s), with all offerings subject to regional variations and governed by local laws, regulations, and company policies. Benefits may vary by country and region, and further details will be provided as part of the recruitment process.

At NetApp, we embrace a hybrid working environment designed to strengthen connection, collaboration, and culture for all employees. This means that most roles will have some level of in-office and/or in-person expectations, which will be shared during the recruitment process.

Equal Opportunity Employer:

NetApp is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination based on age, race, color, gender, sexual orientation, gender identity, national origin, religion, disability or genetic information, pregnancy, protected veteran status, and any other protected classification.

Why You'll Thrive at NetApp

At NetApp, you won't wait for the perfect moment—you'll make it. The early planning, the extra thought, the bold idea that turns good into great: That's how our people operate and how we continue to push the boundaries of data infrastructure.

NetApp is the trusted partner for organizations transforming data into opportunity. As the only enterprise-grade storage service natively embedded in Google Cloud, AWS, and Microsoft Azure, we empower customers to run everything from traditional workloads to enterprise AI with unmatched performance, resilience, and security.

Our culture

We celebrate mold breakers, bold thinkers, and problem solvers. We reward initiative, impact, and ownership. We provide flexibility so you can balance professional ambition with your personal life. Here, differences are not just welcomed—they drive everything we do.

If you're ready to innovate, rise to the challenge, and own every moment - make your next move your best one. Apply now.

Submitting an Application

To ensure a streamlined and fair hiring process for all candidates, our team only reviews applications submitted through our company website. This practice allows us to track, assess, and respond to applicants efficiently. Emailing our employees, recruiters, or Human Resources personnel directly will not influence your application.

AI Disclosure

For select roles, some stages of our hiring process may use artificial intelligence tools to help evaluate applications and candidate selection. These tools support—rather than replace—human decision-making.

Get similar jobs in United States by email

We'll email you when new jobs similar to this one appear.

Similar jobs

Explore more Software Developer jobs in United States.

No similar openings right now. Refine your search or create an alert above.