Site Reliability Engineer - DeviantArt
Wix · Remote, Oregon
Posted Oct 9, 2026 · Verified open Oct 9, 2026
Apply on the employer's siteFind more jobs like this on LandMeAbout the job
- We are seeking a Site Reliability Engineer to architect, scale, and maintain DeviantArt’s high-throughput AWS infrastructure, supporting over 1.5 billion monthly page views. In this role, you will drive infrastructure automation, optimize high-scale database performance, harden system security, and ensure maximum uptime through robust deployment pipelines and proactive site reliability management.
- High-Scale AWS Infrastructure: Architect and maintain auto-scaling AWS systems designed to reliably serve DeviantArt's 1.5B+ monthly page views.
- High Availability & Incident Response: Maximize platform uptime and rapidly resolve performance degradation or service outages.
- Database & Search Optimization: Scale, maintain, and troubleshoot sharded MySQL databases and search infrastructure (e.g., Vespa.ai) for low-latency queries.
- CI/CD & Dev Parity: Build zero-downtime pipelines (Terraform, Kubernetes, GitHub Actions, CodeBuild) and maintain production-parity dev environments.
- Infrastructure Automation: Automate resource provisioning and configuration management, backed by clear documentation and automated testing.
- Security & DDoS Mitigation: Enforce robust security protocols, defend against targeted DDoS attacks, and execute routine patch management.
- Cloud Cost Optimization: Monitor and tune AWS resource utilization to minimize infrastructure spend without compromising site performance.
- 24/7 Operations & On-Call: Ensure continuous reliability as part of a 3-engineer 24/7 on-call rotation.
- Minimum of 10 years of experience working with systems at scale in either a Dev Ops, Platform Engineer, or Site Reliability Engineer type role.
- Excellent analytical skills with the ability to troubleshoot complex problems, analyze system bottlenecks, and implement effective solutions, from frontend through backend systems, sometimes during production degradation or outage.
- Exceptional command line Linux skills.
- In-depth knowledge of AWS services, infrastructure as code using Terraform, and container orchestration with Kubernetes.
- Experience building Docker images, and composing Helm charts.
- Proficiency in Python and Bash for scripting and automation.
- Experience with sharded MySQL databases.
- Proven track record in security compliance standards on large-scale web infrastructures and in-depth knowledge of DDOS mitigation strategies.
- A proactive mindset in identifying potential issues and taking pre-emptive actions to prevent downtime or performance degradation.
- Excellent communication skills, open minded, and capable of effectively collaborating with cross-functional teams and articulating technical concepts to non-technical stakeholders.
- Bonus points: if you have experience with GitOps tools and methodologies.