Platform Engineer
Storm2
⚡ Platform Engineer – Infrastructure & AI Systems
🌍 San Francisco, CA (On-site)
💲 $140,000-$280,000 + Bonus + Equity
The Company
Storm2’s client is a fast-growing insurtech business rebuilding commercial insurance distribution with AI at its core. They’re tackling a market where millions of businesses remain uninsured or underinsured due to outdated processes, slow decision-making, and operational complexity. Having achieved significant growth over the past 18 months, they’re now investing heavily in the infrastructure that underpins the next stage of scale.
The Role
This is a platform engineering role for someone who wants leverage. Not the kind that comes from managing a team, but the kind that comes from building systems that make every other engineer more effective.
The business operates hundreds of services across cloud environments, supports large-scale AI workloads, and processes thousands of automated decisions every day. As the platform grows, reliability, observability, performance, and cost efficiency become critical engineering problems rather than operational concerns.
You’ll work closely with founders and product engineers to build the infrastructure that keeps everything moving. The right person has owned production systems before, understands what breaks at scale, and prefers solving the root cause rather than accepting recurring operational pain.
What you’ll be working on:
Storm2’s client is a fast-growing insurtech business rebuilding commercial insurance distribution with AI at its core. They’re tackling a market where millions of businesses remain uninsured or underinsured due to outdated processes, slow decision-making, and operational complexity. Having achieved significant growth over the past 18 months, they’re now investing heavily in the infrastructure that underpins the next stage of scale.
This is a platform engineering role for someone who wants leverage. Not the kind that comes from managing a team, but the kind that comes from building systems that make every other engineer more effective.
The business operates hundreds of services across cloud environments, supports large-scale AI workloads, and processes thousands of automated decisions every day. As the platform grows, reliability, observability, performance, and cost efficiency become critical engineering problems rather than operational concerns.
- Owning core cloud infrastructure, databases, networking, and CI/CD systems
- Building resilient, scalable platforms capable of supporting high-volume AI workloads
- Designing observability tooling that identifies both technical failures and silent operational issues
- Improving platform reliability through SLOs, incident management, automation, and proactive engineering
- Developing internal tooling and infrastructure that accelerates developer productivity
- Orchestrating large numbers of concurrent, non-deterministic AI workflows
- Building evaluation frameworks and systems that measure the quality and performance of AI outputs
- Managing infrastructure performance and cloud cost optimisation as usage scales
What you’ll bring:
- Experience in Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering
- Strong software engineering background with Python, TypeScript, Go, or similar languages
- Production experience with AWS and cloud-native infrastructure
- Deep understanding of at least two of the following: databases, networking, observability, cloud infrastructure, or CI/CD
- Experience designing and operating distributed systems under meaningful production load
- A track record of building platforms, tooling, or services that other engineers rely on
- Strong operational mindset with experience owning reliability and production incidents
- Familiarity with modern AI-assisted development tools and workflows
Nice to have:
- Experience running AI, ML, or LLM-powered workloads in production environments
- Knowledge of real-time systems or voice technologies
- Deep AWS expertise across services such as ECS, Lambda, RDS, and VPC architecture
- Startup experience within high-growth or scaling environments
