ML Data & Platform Engineer at Speechmatics
- Location
- London, United Kingdom
- Compensation
- Not Disclosed
Speechmatics is searching for an ML Data & Platform Engineer to lead a critical greenfield mandate building a modern ML platform from the ground up. You will tackle the massive data scale challenges unique to speech AI, bridging the gap between data engineering and production ML in a high-impact, strategic seat.
This isn't a maintenance role; you will own the full data and infrastructure layer powering frontier-level models for global enterprise customers. If you are a technical builder with experience in Python, distributed systems, and GPU optimization, join this London-based team and help us achieve our mission to Understand Every Voice.
Role overview
You will own the full data and infrastructure layer behind Speechmatics’ world-class speech models. This greenfield mandate involves building a modern ML platform from scratch to replace legacy systems while driving a complex ML data strategy. You’ll bridge the gap between data engineering and production ML to accelerate model delivery.
About Speechmatics
Software
Speechmatics is a technology company specializing in artificial intelligence-driven speech recognition and voice AI infrastructure. It provides automatic speech recognition (ASR), speech-to-text (STT), text-to-speech (TTS), translation, summarization, sentiment analysis, and topic detection services, supporting over 50 languages and dialects with high accuracy across accents, noisy environments, and multi-speaker scenarios. The company serves enterprises, developers, and partners in industries such as media & broadcast, contact centers, healthcare, education, legal, finance, and AI infrastructure, enabling real-time and batch transcription for applications like captioning, compliance, and voice agents.[1][5]
What you will do
- Design and build scalable data pipelines in Python and SQL for ingesting and transforming massive datasets, including automated web scraping and acquisition solutions.
- Develop a high-performance ML platform to train, evaluate, and serve models, optimizing for GPU utilization, job scheduling, and distributed training efficiency.
- Implement production-grade observability and troubleshooting for distributed systems to ensure the reliability and speed of the entire ML lifecycle.
Who this is a fit for
- Strong background in backend or data engineering with deep proficiency in Python, SQL, and hands-on experience with Docker, Kubernetes, and major cloud providers.
- Proven experience building large-scale ETL/ELT pipelines and automated data collection systems, ideally within a high-growth speech AI or frontier model environment.
- A builder’s mindset with a solid understanding of the full ML lifecycle, from data curation and MLOps to model serving and distributed training optimization.
Why this role is remarkable
- True Greenfield Mandate: You aren’t maintaining a legacy stack; you’ll build a modern, SaaS-based ML platform from the ground up to replace existing internal systems.
- Unique Scale Challenges: Solve the specialized data problems unique to Speech AI, managing the sourcing, validation, and storage of massive training datasets for frontier-level models.
- Strategic Ownership: Work as a core part of the ML team with the autonomy to identify friction and define the technical roadmap, rather than just executing a backlog.
How Jack & Jill work together

What happens next?
Jack’s an AI agent for job searching and career coaching. He works for you.
Jill is the AI recruiter working for the company. She recruits from Jack’s network.
If your profile’s a match and Speechmatics wants to meet, Jill will make the intro. In the meantime, Jack will send you excellent alternatives.