Data Engineer
Remote
Full Time
Mid Level
Hypersonix.ai is disrupting the e-commerce space with AI, ML and advanced decision capabilities to drive real-time business insights. Hypersonix.ai has been built ground up with new age technology to simplify the consumption of data for our customers in various industry verticals.
We are looking for a skilled and motivated Data Engineer to design, build, and maintain scalable data pipelines and data infrastructure. You will work closely with Data Science, Engineering, and Product teams to ensure high-quality, reliable, and accessible data for analytics, machine learning, and business applications.Roles and Responsibilities
- Design, develop, and maintain scalable data pipelines for batch and real-time data processing.
- Build and optimize ETL/ELT workflows to collect, transform, and load data from multiple sources.
- Develop and maintain data models, data warehouses, and data lakes.
- Ensure data quality, consistency, accuracy, and reliability across data pipelines.
- Optimize data processing workflows for performance, scalability, and cost efficiency.
- Work with large and complex datasets from multiple sources.
- Collaborate with Data Scientists and ML Engineers to prepare and deliver high-quality datasets for machine learning models.
- Develop data integrations with APIs, databases, cloud platforms, and third-party systems.
- Monitor data pipelines and troubleshoot data processing and infrastructure issues.
- Implement data validation, monitoring, and testing frameworks.
- Maintain clear documentation of data pipelines, data models, and technical processes.
- Follow best practices for data security, governance, and access control.
- Contribute to the design and evolution of Hypersonix's data platform and architecture.
- 5–7 years of experience in Data Engineering or a similar role.
- Strong programming skills in Python or another programming language.
- Strong experience with SQL and relational databases.
- Hands-on experience building ETL/ELT pipelines and data workflows.
- Experience with data processing technologies such as Spark/PySpark.
- Experience working with cloud platforms such as AWS, Azure, or GCP.
- Good understanding of data warehousing and data lake concepts.
- Experience with workflow orchestration tools such as Airflow, Dagster, or similar.
- Experience working with distributed systems and large-scale datasets.
- Strong understanding of data modeling, database design, and performance optimization.
- Familiarity with Git and CI/CD practices.
- Good understanding of software engineering principles and coding best practices
Apply for this position
Required*