Careers at BashClouds

Site Reliability Engineer

Make system behaviour visible before it becomes an incident.

OPEN AREA

Observability, SLOs, and incident response for production systems.

Role overview

Make system behaviour visible before it becomes an incident.

You would design observability that supports day-to-day operations, not just more dashboards, and connect technical signals to decisions people can act on.

This is a general expression of interest for this area, not a promise of a currently open, funded position. Tell us what you are good at, what you want to learn, and which of these problems you enjoy solving.

Email hr@bashclouds.com

What you would work on

  • Design metrics, logs, traces, and event strategy
  • Review alerts and reduce noise
  • Define service-level indicators and objectives
  • Support incident response and post-incident learning
  • Review capacity, performance, and reliability

Who tends to fit

Practical experience matters more than titles.

You have carried a pager and want the next incident to be less confusing than the last one, for the whole team, not just for you.

Useful experience

  • Production on-call or incident response experience
  • Metrics, logging, and tracing tooling
  • SLIs, SLOs, and error-budget practices
  • Alert design and noise reduction
  • Capacity and performance analysis
  • Post-incident review facilitation