Staff Software Engineer, Data Catalog
About the role
About Our Client:
This organization operates in the online hospitality and travel industry, connecting hosts and guests worldwide. It addresses the challenge of enabling authentic community connections through unique stays and experiences. With millions of hosts and billions of guest arrivals globally, the organization has a significant reach across nearly every country.
About the Opportunity:
The Staff Software Engineer, Data Catalog will lead the development of data catalog infrastructure within the Data Infrastructure team. This role focuses on creating tools and processes to ensure data assets are cataloged, discoverable, and governed effectively. The position contributes to improving data management and governance across the organization's data ecosystem.
Responsibilities:
Develop and implement data catalog products to support data discovery, metadata management, data quality, and data lineage.
Build metadata infrastructure to support data governance, including ownership management, classification, privacy, access control, and retention.
Create comprehensive data lineage products covering various data types and sources.
Collaborate with cross-functional teams to catalog and govern datasets in line with business goals and regulatory requirements.
Design and build metadata integrations with data infrastructure frameworks.
Develop tools and processes for scalable metadata onboarding and integration.
Implement metadata-driven data policies and procedures supporting governance needs.
Requirements:
BS/MS/PhD in Computer Science or related field, or equivalent experience.
Over 9 years of software engineering experience focused on data infrastructure.
Experience with data storage and distributed processing technologies such as Hive, Spark, Trino, Flink, or SQL databases.
Proficient in Java, Python, or Scala programming languages.
Experience with data catalog products and metadata management frameworks.
Knowledge of workflow orchestration tools like Apache Airflow, Prefect, or Kubeflow.
Strong understanding of data governance frameworks including classification, lineage, quality, privacy, and retention.
Excellent communication skills for cross-team collaboration.
Ability to work beyond usual areas of expertise.
Strong analytical and problem-solving skills.
Pay Range and Compensation Package:
Pay range is $212,000 to $265,000 USD, with potential eligibility for bonus, equity, benefits, and employee travel credits.
Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.
Note:
RemoteHunter is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.
Responsibilities
- Develop and implement data catalog products to support data discovery, metadata management, data quality, and data lineage.
- Build metadata infrastructure to support data governance, including ownership management, classification, privacy, access control, and retention.
- Create comprehensive data lineage products covering various data types and sources.
- Collaborate with cross-functional teams to catalog and govern datasets in line with business goals and regulatory requirements.
- Design and build metadata integrations with data infrastructure frameworks.
- Develop tools and processes for scalable metadata onboarding and integration.
- Implement metadata-driven data policies and procedures supporting governance needs.
Qualifications
- BS/MS/PhD in Computer Science or related field, or equivalent experience.
- Over 9 years of software engineering experience focused on data infrastructure.
- Experience with data storage and distributed processing technologies such as Hive, Spark, Trino, Flink, or SQL databases.
- Proficient in Java, Python, or Scala programming languages.
- Experience with data catalog products and metadata management frameworks.
- Knowledge of workflow orchestration tools like Apache Airflow, Prefect, or Kubeflow.
- Strong understanding of data governance frameworks including classification, lineage, quality, privacy, and retention.
- Excellent communication skills for cross-team collaboration.
Benefits
- Potential eligibility for bonus
- Equity
- Benefits
- Employee travel credits
Skills mentioned
About RemoteHunter
RemoteHunter is your dedicated AI job search assistant, turning the job hunt from a slow, individual effort into a quicker, smarter, and guided experience by streamlining each step of the process and speeding up your path to the right career opportunities.