Job DescriptionThis role provides support to end users and responds to incident and problem management issues for multiple applications, focusing on leading triage activities on all business affecting incidents. Ensures compliance with incident and problem management policies and procedures. Serves as a key focal point for the customer, client, and associate experience, restoring any impacts regardless of root cause.ResponsibilitiesLeads production support triage efforts, manages bridge line troubleshooting, engages in technical research, and escalates issues to leadership as neededEnsures all impacts are accurately recorded and documented in the system of record, verifies documents and wikis are updated and available for use during triage, and supports on call responsibilities for incidents, the documentation of application flows, impacts during outages, the customer experience, and contacts for support needsProvides status updates and technical detail for awareness communications, such as infrastructure, application and client impact, and component points of failure, oversees accuracy of all communications sent, and ensures any necessary reconvenes are scheduledIdentifies business impact, interprets monitors, dashboards, and logs, and writes queries to accurately calculate and communicate impacts to leadership in partnership with senior team members or specialists within Technology ServicesPromotes and enforces production governance during triage/testing, and identifies production failure scenarios, vulnerabilities, and opportunities for improvement, determines appropriate actions, and escalates issues as neededAnalyzes, manages, and coordinates incident management activities to detect problems that potentially affect the service levelFulfills research requests, ad hoc reports, and offline incidents at the direction of senior team members or the Technology/Production Services teamsRequired QualificationsHands‑on experience with Splunk (search, SPL, dashboards, alerts, data onboarding, and tuning).Hands‑on experience with Dynatrace (APM, services/entities, alerting profiles, management zones, dashboards).Strong understanding of monitoring and observability concepts: logs, metrics, traces, events, and correlation.Experience supporting production systems and participating in incident management and operational support.Knowledge of SRE concepts such as reliability engineering, alert hygiene, post‑incident reviews, and automation.Experience working with ITSM processes (incident, problem, change) and tracking SI actions to closure.Basic to intermediate scripting experience (e.g., Python, Shell) for automation and analysis.Strong communication skills and ability to work across distributed teams in the APAC region.Desired QualificationsExperience with advanced Splunk or Dynatrace features (custom metrics, anomaly detection, DQL/SPL optimization, synthetic monitoring).Experience integrating monitoring tools with ServiceNow or similar ITSM platforms.Familiarity with capacity monitoring, performance engineering, or business transaction monitoring.Relevant certifications (Splunk, Dynatrace, SRE/DevOps, Cloud) are a plus.SkillsAdaptabilityAnalytical ThinkingInfluenceProduction SupportRisk ManagementAutomationCollaborationInnovative ThinkingResult OrientationSolution DesignBusiness AcumenDevOps PracticesProject ManagementSolution Delivery ProcessStakeholder ManagementShift1st shift (United States of America)Hours Per Week40#J-18808-Ljbffr