Dashboards can lie by omission. A service might look perfectly healthy in aggregate while quietly failing for one specific subset of users, and the difference between catching that in ten minutes versus two hours usually comes down to how good the observability tooling actually is. That is what this full-time, fully remote engineering role is responsible for building and maintaining.
Observability as a discipline covers metrics, logs, and traces together, rather than any one of them in isolation. A metric might show latency creeping up, but only a trace can show which specific downstream service is actually responsible, and only good logging can explain why. Building tooling that connects those three views cleanly is most of what separates useful observability from a wall of dashboards nobody trusts.
A meaningful part of the role involves being the person other engineers come to when something is slow or broken and nobody can immediately say why. That means the tooling you build has to hold up under pressure, not just look good in a demo, since the whole point of good observability is that it keeps working when everything else is on fire.
Work usually splits between longer-term projects, like improving how traces connect across services, and shorter reactive tasks tied to whatever incident or alert came up that week. Code review and documentation tend to fill in the gaps between the two, and a good week often ends with slightly better visibility into the system than it started with, even if no single project felt finished.
Incident response weaves through both categories. A production alert can pull attention away from a planned project for an afternoon, and part of the role is treating that interruption as useful signal about what the tooling still misses, rather than just a distraction to get through as quickly as possible.
A quarter might involve shipping a new distributed tracing pipeline in the first half, then spending several weeks in the second half responding to how that pipeline actually performs once real traffic hits it. The two phases feed each other directly, which is part of what makes the role feel connected end to end rather than handed off between separate teams.
Engineers who enjoy seeing the direct results of their own tooling in a live incident tend to find that loop satisfying rather than exhausting. Others find the constant feedback draining, and it is worth being honest with yourself about which camp you fall into before taking the role, since the pace here rarely lets up for long stretches at a time.
Candidates typically hold a bachelor's degree in computer science, software engineering, or a related field, along with demonstrated experience in observability specifically. A strong portfolio of relevant projects matters, along with proficiency across modern development tools. 24 months of hands-on experience is required, and most successful candidates have worked directly with production systems where downtime or slow performance had a real, measurable cost.
A candidate who has previously instrumented a service from scratch, adding metrics and tracing where none existed before, usually has a better feel for this role than one who has only consumed dashboards someone else built. Knowing what data is actually worth capturing, instead of instrumenting everything and drowning the team in noise, is a skill that takes real hands-on time to develop.
Specific tool experience helps but matters less than the underlying practice, since teams standardize on different stacks. What carries across any stack is the habit of asking what a given signal would actually tell an on-call engineer at 3 a.m., before adding it to a dashboard nobody has time to parse under pressure.
This is a fully remote, worldwide position with no office tied to it. Because observability work often intersects with incident response, some overlap with the wider engineering team is expected so that alerts and escalations do not sit unanswered for hours. Most day-to-day collaboration runs through code review, written design docs, and asynchronous updates in a shared chat tool, with synchronous meetings kept for the discussions that genuinely need real-time back-and-forth. Remoteroles works with employers who structure on-call rotations fairly across time zones rather than defaulting to one region, which matters a great deal for a role where a 2 a.m. page is always a possibility.
This role pays $143,000 a year, full-time. Full-time observability roles at this level commonly include:
Given the seniority of this position, candidates should expect the compensation conversation to include specifics on stipend amounts and any additional equity or bonus structure during the interview process.
Compensation at this level generally reflects how much production risk the role carries. An engineer whose tooling directly shortens incident response time across an entire platform is compensated closer to a senior software engineer than to a support-focused operations role, which is part of why this position sits well above average for remote engineering work in general.
Submit your resume along with links to relevant projects or systems you have worked on, particularly anything showing observability tooling you built or improved. Interviews typically include a technical discussion of a past incident you helped diagnose, so it helps to have a specific example ready rather than a general summary of your experience.