Fusemachines
10 months ago
Lead Data Engineer
Sign up to save this job, get alerts, and apply with an optimized CV.
Company information
- Company
- Fusemachines
- Location
- Canada, Ontario, Toronto Canada
- Posted
- 10 months ago
Job description
About Fusemachines Fusemachines is a leading AI strategy, talent, and education services provider. Founded by Sameer Maskey Ph.D., Adjunct Associate Professor at Columbia University, Fusemachines has a core mission of democratizing AI. With a presence in 4 countries (Nepal, United States, Canada, and Dominican Republic and more than 400 full-time employees). Fusemachines seeks to bring its global expertise in AI to transform companies around the world. About the Role Location: Remote (Full-time) We are looking for a Lead Data Engineer to join our growing team. As a Lead Data Engineer, you will be responsible for designing, developing, and maintaining our data infrastructure. You will work closely with other engineers, data scientists, and business stakeholders to understand data needs, build data pipelines, and ensure data quality and availability. Responsibilities Design, build, and maintain scalable and reliable data pipelines using tools like Apache Spark, Airflow, and cloud-based data platforms. Collaborate with data scientists and other engineers to understand data requirements and build data solutions that meet their needs. Develop and maintain data models, schemas, and data governance policies. Monitor data pipelines and infrastructure for performance, availability, and cost optimization. Implement data quality checks and ensure data accuracy and consistency. Lead and mentor a team of data engineers, providing technical guidance and support. Stay up-to-date with the latest technologies and trends in data engineering. Requirements Bachelor's or Master's degree in Computer Science, Engineering, or a related field. 7+ years of experience in data engineering, with a focus on building and maintaining data pipelines. Strong experience with data warehousing, ETL processes, and data modeling. Proficiency in programming languages such as Python or Scala. Experience with big data technologies such as Apache Spark, Hadoop, and Hive. Experience with cloud-based data platforms such as AWS (e.g., S3, EMR, Redshift), Azure (e.g., Azure Data Lake Storage, Azure Databricks), or Google Cloud Platform (e.g., Google Cloud Storage, Dataproc). Experience with data pipeline orchestration tools such as Airflow or Luigi. Experience with data governance and data quality best practices. Strong communication and collaboration skills. Experience with Agile development methodologies. Nice to Have Experience with real-time data streaming technologies such as Kafka or Spark Streaming. Experience with NoSQL databases such as Cassandra or MongoDB. Experience with data visualization tools such as Tableau or Power BI. Benefits Competitive salary and benefits package. Opportunity to work on challenging and impactful projects. Collaborative and supportive work environment. Opportunity for professional growth and development. Remote work flexibility. Equal Opportunity Employer Fusemachines is an Equal Opportunity Employer that values diversity at all levels. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability. Please review our privacy policy here: https://www.fusemachines.com/privacy-policy/
Required skills
- kafka
- power bi
- etl
- data modeling
- python
- aws
- agile
- s3
- airflow
- hadoop
- hive
- azure
- cassandra
- google cloud platform
- nosql
- data governance
- mongodb
- tableau
- emr
- data engineering
- data pipelines
- scala
- data warehousing
- redshift
- dataproc
- azure data lake storage
- luigi
- apache spark
- azure databricks
- spark streaming
- google cloud storage
Interested in this position?
Create your free account and tailor your CV to match this job.