Applications are completed on the employer's own site.
About this role
Senior Software Engineer, Compute Platform
$130k–290k🌍 Remote-friendly📍 Foster City, CA
AT A GLANCE
Senior distributed systems engineer to build and optimize Replit's cloud infrastructure platform for application deployment and scaling.
LEVEL
senior
KEY SKILLS
GolangRustDistributed systemsCloud infrastructureKubernetesGoogle Cloud PlatformContainer orchestrationServerless computingLinux internalsSystem optimizationMonitoring and alertingAuto-scalingCost optimizationPlatform-as-a-service
WHAT YOU'LL DO
- Expand Replit's cloud infrastructure offerings and launch new cloud products for the Replit Agent
- Enhance reliability and scalability by identifying bottlenecks and implementing monitoring systems
- Optimize cloud infrastructure utilization and reduce costs through resource provisioning and auto-scaling strategies
- Collaborate with cross-functional teams and SRE to ensure high availability and minimal downtime
- Participate in on-call rotation for infrastructure support
THE ROLE
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.
About the role:
We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you.
In this role you will
- Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently.
- Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base.
- Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services.
Required skills and experience:
- Distributed systems: Track record of working with platform-as-a-service, distributed storage, or information retrieval systems. Experience in designing scalable architectures and optimizing systems for latency or cost.
- Problem-solving mindset: Ability to approach complex challenges pragmatically and devise effective solutions. You think radically but ship incrementally.
- Self-directed and autonomous: Able to work independently, set priorities, and drive projects forward. You take ownership and initiative.
- Versatility and flexibility: Able to wear multiple hats and tackle a wide range of challenges. You are comfortable working across different layers of the stack and adapting to the needs of the project.
- Continuous learning and adaptability: Passionate about staying up-to-date with industry trends and expanding your skill set. You embrace change and adapt quickly.
Nice to have:
- Experience working on cloud infrastructure or platform products, particularly in the areas of application deployment, serverless computing, or container orchestration.
- Familiarity with Google Cloud Platform (GCP) services and tools, such as GCE, GKE,, Cloud Run, or Cloud Storage.
- Contributions to open-source projects related to cloud technologies, deployment frameworks, or developer tools. We love OSS!
Tools + Tech Stack for this role
This role may not be a fit if
- You are a generalist backend engineer who hasn’t built scalable distributed systems.
- You cannot take part in the oncall rotation of min 6 people.
- You do not enjoy diving into Linux internals.
This is a full-time role that can be held from our Foster City, CA office. The hybrid role has an in-office requirement of Monday, Wednesday, and Friday.
Full-Time Employee Benefits Include
💰 Competitive Salary & Equity
💹 401(k) Program with a 4% match (US Only)
⚕️ Health, Dental, Vision and Life Insurance
🩼 Short Term and Long Term Disability
🚼 Paid Parental, Medical, Caregiver Leave
🏝 Flexible Time Off (FTO) + Holidays
🚗 Commuter Benefits (In-Office & US Only)
📱 Monthly Wellness Stipend
🧑💻 Autonomous Work Environment
🖥 In Office Set-Up Reimbursement (In-Office Only)
🚀 Quarterly Team Gatherings
☕ In Office Amenities (In-Office Only)