Search

Senior Reliability Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Job Description

We are looking for a Senior Reliability Engineer to strengthen the stability, scalability, and performance of technology systems that support construction and field operations. This position plays a key role in building dependable infrastructure across remote job sites, field offices, and cloud environments so teams can work efficiently with minimal disruption. The ideal candidate brings a strong SRE mindset, combines technical depth with practical problem-solving, and partners effectively with operational teams to maintain business continuity and system resilience.


Responsibilities:

• Build and enhance highly available infrastructure that supports office locations, remote field environments, networking needs, cloud services, and edge-based systems.

• Direct incident response efforts for service disruptions, coordinate restoration activities, and lead root cause investigations to prevent repeat issues.

• Create and maintain monitoring, alerting, and observability capabilities that improve visibility into system health, uptime, and application performance.

• Work closely with construction, engineering, and field personnel to ensure technology reliability aligns with project schedules, operational demands, and safety expectations.

• Implement automated infrastructure deployment and recovery processes using Infrastructure as Code and configuration management tools such as Terraform and Ansible.

• Establish service reliability targets, manage service level objectives, and use error budgets to guide operational decisions and continuous improvement.

• Strengthen the security posture of remote and field-deployed systems by improving hardening practices and secure access methods.

• Provide guidance to less experienced engineers and help foster a culture centered on reliability, accountability, and operational excellence.

• Identify process improvements that reduce inefficiencies, simplify support efforts, and improve overall service delivery.


• Relevant experience in site reliability, infrastructure engineering, or a closely related IT operations role.
• Strong understanding of cloud platforms, enterprise networking, and edge computing concepts in distributed environments.
• Hands-on experience with Infrastructure as Code, including tools such as Terraform and Ansible.
• Proficiency with monitoring and observability platforms such as Splunk and Dynatrace, along with scripting or automation capabilities.
• Familiarity with construction, industrial, or field-based technology environments is strongly preferred.
• Knowledge of systems used in operational technology settings, including SCADA, IoT-connected devices, or similar enterprise platforms.
• Excellent written and verbal communication skills with the ability to collaborate effectively across technical and operational teams.
• Ability to handle sensitive information with discretion while working independently and contributing positively in a team setting.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...