RITS
1 month ago
Data Engineer
Sign up to save this job, get alerts, and apply with an optimized CV.
Company information
- Company
- RITS
- Location
- Polska, mazowieckie, Warszawa Poland
- Posted
- 1 month ago
Job description
We are looking for a skilled Data Engineer to join our team and help build, develop, and operate our next-generation data platform. In this role, you will work with modern cloud technologies, data engineering tools, and machine learning pipelines to deliver reliable, scalable, and mission-critical data solutions. Job Responsibilities: • Build and run Tradeweb’s data platform using such technologies as public cloud infrastructure (AWS and GCP), Kafka, Spark, databases and containers • Develop Tradeweb’s data platform based on open source software and Cloud services • Build and run ETL pipelines to onboard data into the platform, define schema, build DAG processing pipelines and monitor data quality. • Help develop machine learning development framework and pipelines • Manage and run mission crucial production services. Key Details: Work Model: 100% Remote Working Hours: NYC Time Zone (Minimum 5h overlap required) Contract & Rate: B2B, 40 - 60 USD Requirements: • Strong eye for detail, data precision, and data quality. • Strong experience maintaining system stability and responsibly managing releases. • Considerable production operations and support experience. • Clear and effective communicator who is able to liaise with team members and end-users on requirements and issues. • Agile, self-starter who is able to responsibly see things through to completion with minimal assistance and oversight. • Expert level grasp of SQL and databases/persistence technologies such as MySQL, PostgreSQL, SQL Server, Snowflake, Redis, Presto, etc • Strong grasp of Python and related ecosystems such as conda or pip. • Experience building ETL and stream processing pipelines using Kafka, Spark, Flink, Airflow/Prefect, etc • Experience with using AWS/GCP (S3/GCS, EC2/GCE, IAM, etc), Kubernetes and Linux in production. • Experience with parallel and distributed computing • Strong proclivity for automation and DevOps practices and tools such as Gitlab, Terraform, Prometheus. • Experience with managing increasing data volume, velocity and variety. • Ability to deal with ambiguity in a changing environment. • At least 5-6 hours overlap starting from 9am US Eastern. Good to have: • Familiarity with data science stack: e.g. Jupyter, Pandas, Scikit-learn, Pytorch, MLFlow, Kubeflow etc • Development skills in Java, Go, or Javascript • Software builds and packaging on MS Windows • Experience managing time series data • Familiarity with working with open source communities • Financial Services experience
Required skills
- mysql
- redis
- kafka
- gitlab
- snowflake
- remote
- java
- databases
- javascript
- financial services
- sql
- python
- aws
- kubernetes
- pytorch
- automation
- ec2
- data engineer
- airflow
- presto
- spark
- cloud technologies
- gcp
- containers
- postgresql
- devops
- go
- terraform
- prometheus
- machine learning
- linux
- sql server
- data quality
- b2b
- distributed computing
- pip
- data platform
- iam
- ms windows
- mlflow
- etl pipelines
- pandas
- scikit-learn
- kubeflow
- jupyter
- parallel computing
- prefect
- flink
- gcs
- time series data
- open source communities
- gce
- open source software
- conda
- machine learning pipelines
- production services
- data engineering tools
- data volume
- dag processing
- nyc time zone
- data velocity
- data variety
Interested in this position?
Create your free account and tailor your CV to match this job.