Our client is a technology-oriented company developing advanced software and artificial intelligence systems to address complex, high-value problems. The organization is building a modern technology platform that combines scalable software infrastructure, data-intensive workflows, and machine learning capabilities.
Its activities rely on close collaboration between software and machine learning engineers, with a strong focus on engineering quality, reliability, and development velocity. As an early-stage organization, the company offers the opportunity to shape foundational architecture, build critical platform capabilities from the ground up, and create the infrastructure required to support increasingly sophisticated AI and agentic systems at scale.
Mission
The Senior Software Engineer will be responsible for designing and building the core software platform and infrastructure supporting the company's computational, data, and machine learning workflows. The role will combine backend engineering, distributed systems, data infrastructure, and production reliability, with particular ownership of the foundations required to deploy and operate AI-driven and agentic systems across complex scientific workflows.
Responsibilities
Backend and Platform Engineering
Design, build, and maintain backend services and platform capabilities supporting scientific applications, machine learning systems, and internal research workflows.
Develop robust APIs, services, and abstractions that enable researchers and engineers to interact efficiently with data, models, computational resources, and internal tooling.
Make pragmatic architectural decisions balancing development velocity, system simplicity, scalability, and long-term maintainability.
Data and Computational Workflows
Build reliable data pipelines and workflow orchestration systems for heterogeneous scientific and machine learning workloads.
Develop infrastructure for moving, transforming, validating, and tracking data across computational and experimental workflows.
Improve reproducibility, observability, and execution reliability across long-running and compute-intensive processes.
Reliability and Infrastructure
Own the reliability, performance, monitoring, and operational quality of critical platform services and infrastructure.
Design systems capable of scaling with increasing data volumes, computational workloads, model complexity, and organizational growth.
Establish engineering practices around testing, observability, deployment, incident prevention, and infrastructure automation.
AI and Agentic Systems
Build the software infrastructure required to integrate models, tools, data sources, and computational environments into agentic workflows.
Develop reliable orchestration and execution layers enabling AI systems to perform multi-step tasks across scientific and engineering environments.
Address practical challenges around state, tool execution, failure recovery, evaluation, traceability, and human interaction within production-grade agentic systems.
Required Qualifications
Strong professional experience in backend, platform, infrastructure, or distributed systems engineering.
Excellent software engineering skills, with strong proficiency in Python and experience building production-grade backend systems.
Demonstrated experience designing APIs, services, data pipelines, or workflow orchestration infrastructure.
Strong understanding of distributed systems, databases, cloud infrastructure, and modern software architecture.
Experience building reliable production systems with appropriate testing, monitoring, observability, and operational practices.
Ability to independently own complex technical problems from architecture and implementation through deployment and operation.
High degree of ownership and comfort operating within an early-stage, fast-moving, technically demanding environment.
Strong collaboration and communication skills across engineering, machine learning, and scientific teams.
Preferred Experience
Experience building platforms or infrastructure for machine learning, scientific computing, computational biology, or other data-intensive technical environments.
Experience with workflow orchestration, distributed compute, containerization, and cloud-native infrastructure.
Familiarity with LLM-based applications, AI agents, tool-use frameworks, or production agentic systems.
Experience designing infrastructure for asynchronous, long-running, or failure-prone computational workflows.