About impact.com impact.com is the world’s leading commerce partnership marketing platform transforming the way businesses grow by enabling them to discover manage and scale partnerships across the entire customer journey. From affiliates and influencers to content publishers brand ambassadors and customer advocates impact.com empowers brands to drive trusted performance-based growth through authentic relationships. Its award-winning products— Performance (affiliate) Creator (influencer) and Advocate (customer referral)—unify every type of partner into one integrated platform. As consumers increasingly rely on recommendations from people and communities they trust impact.com helps brands show up where it matters most. Today over 5000 global brands including Walmart Uber Shopify Lenovo L’Oréal and Fanatics rely on impact.com to power more than 225000 partnerships that deliver measurable business results. Your Role at impact.com As the Site Reliability Engineer for the Content Intelligence & Regulatory Apps Group you will own and grow our reliability practice for the systems that index monitor and enrich social and web content on the impact.com platform. This role focuses on building and professionalizing rather than firefighting. Since the platform is stable your mission is to implement SRE engineering disciplines (service-level objectives observability runbooks and structured root-cause analysis) to ensure reliability is measurable repeatable and owned. You will work with Squad leads Platform Engineering and Cloud Operations and will own the SLO and RCA practice for the group while partnering closely with other internal and external squads. Your job is to provide them with the framework tooling and habits needed to run reliable services and to act as the central point of contact for reliability across those teams. This is a software-engineering-led SRE role where you will read and write Java instrument Spring services tune the JVM help harden infrastructure batch processes data and orchestration flows. Our guiding principle is to prioritize system stability and data integrity above all else. Because these are business-critical systems security and compliance are part of the reliability mandate not an afterthought you will build observability audit trails and operational practices that are secure and auditable by default working alongside the central security and DevOps teams. Reliability resilience data integrity and compliance take precedence over short-term feature velocity. This is a strong opportunity for an engineer ready to step into ownership and grow the role and themselves over time with the support of the group. What You'll Do Own the SLO/SLI practice for CIRA Define meaningful service-level objectives and indicators for services on GCP with each squad. Establish error budgets and necessary baselines applying extra rigor to flows that affect critical business processes. Build in security and compliance by default. Treat security and auditability as reliability properties ensure critical business processes and data flows have the necessary audit trails and if applicable forensic-replay history needed for transactional-correctness and compliance obligations (e.g. SOX and PCI-adjacent concerns). Champion secrets hygiene (HashiCorp Vault GCP Secret Manager SOPS) least-privilege access to production and data and audited break-glass procedures. Partner with the Cloud Security and Cloud Platform teams rather than duplicating their function. Manage vulnerability and patch posture for services Track and drive remediation of vulnerabilities across the JVM Spring/Spring Boot dependencies and container images help establish patching expectations and surface security-relevant findings from quality gates (SonarQube) and secret scanning (ggshield) so they get prioritized alongside reliability work. Own and run root-cause analysis Establish a consistent blameless RCA practice for the group. Drive investigations toward durable fixes and preventative actions while partnering with the owning squad rather than working in isolation. Over time improve the squads' own troubleshooting and post-incident habits. Build and mature observability for critical services Become the group's point person for the observability stack (e.g. Grafana). Build and standardize monitoring dashboards tracing and alerting using Open Telemetry that surface the health of systems services infrastructure and business/transactional processes — without leaking PII data into logs traces or dashboards — while creating reusable patterns that squads can adopt. Keep performance and resource metrics within thresholds Track response latency JVM heap/GC behavior thread-pool saturation CPU/memory consumption cloud costs error rates and uptime across all services (e.g. Java Spring Boot NodeJs etc). Turn findings into prioritized improvements in collaboration with the squads. Improve batch and orchestration reliability Help harden batch jobs and support the migration toward durable workflow orchestration to reduce partial-failure windows and manual idempotency. Support database performance and query optimization Monitor the groups databases (e.g. MySql SingleStore Elastic GCP Spanner) for slow queries indexing lock contention and connection-pool health. Flag operational or performance risks in Liquibase migrations before they ship. Understand client-generated workloads Learn how brand agency and partner workloads translate into resource consumption. Help confirm that consumption aligns with contract tiers and surface anomalous traffic patterns as both a reliability and a security signal escalating suspected abuse to the security team. Establish alerting and runbook standards Define the group's conventions for actionable alerts and runbooks to ensure on-call engineers across squads can respond quickly and consistently. Troubleshoot across the stack Address issues in the JVM/application Spring container MySQL Pub/Sub messaging Feign/REST and legacy RMI/HTTP-Invoker integrations GKE/GCE network and client-generated workloads. This requires growing depth in JVM and database performance (profiling and query optimization). Remediation may include code/query optimizations JVM/connection-pool tuning rate-limit or autoscaling configuration retry/backoff changes or escalating heavy-request patterns. Make the delivery path safer Partner with squads to improve CI/CD (Jenkins/Github Actions) and GitOps deploys (ArgoCD + Helm) including health checks readiness/liveness probes progressive rollouts and rollbacks for finance services. Inform capacity and cost Help analyze platform usage to attribute costs to customers and workloads inform capacity planning and identify efficiency improvements. Partner with Cloud Operations (FinOps). Leave things more reliable than you found them Improve tests observability and operability each time a service is touched and grow the group's reliability guardrails into a shared durable practice. General Other duties as assigned by the Company. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. What You Bring Experience 3+ years in SRE software engineering or systems/operations roles supporting production services with the appetite to own and grow a reliability practice. Systems Design Solid understanding of systems and application design with the ability to reason about reliability failure modes and performance. Java Able to read debug and make changes to Java/Spring code today with the willingness and aptitude to deepen JVM expertise (garbage collection memory thread pools) on the job. Cloud & Kubernetes Experience operating services on a major cloud (ideally GCP/GKE) and exposure to Kubernetes and containers. Observability Proficiency with metrics logging tracing/APM and alerting. Database Experience with relational database monitoring and SQL tuning (MySQL preferred) including basic indexing and query optimization. Automation Proficiency in shell scripting and comfort automating routine operational tasks. Security & Compliance Awareness Practical understanding of operational security for production systems including secrets management least-privilege access and handling sensitive data safely in logs and telemetry. SLO/SLI Familiarity with defining or operating against SLOs/SLIs or a strong desire and aptitude to build that practice. Collaboration A collaborative working style with the ability to influence and enable other engineers. Problem Solving Ability to prioritize work independently and focus on simple efficient and reliable solutions. Education B.S. in Computer Science or a related field or equivalent practical experience. Preferred Qualifications Spring Ecosystem Experience with the Spring/Spring Boot ecosystem and JVM application servers. Batch/Workflow Orchestration Experience with tools like Quartz Temporal or similar. Event-Driven Messaging Familiarity with Pub/Sub or Kafka and service-to-service integration (REST/Feign gRPC). Telemetry & Metrics Analysis Experience with Prometheus/PromQL Grafana and log analytics. AI-assisted engineering Familiarity with using AI tools to assist in engineering workflows. Tooling Familiarity with secrets management (Vault GCP Secret Manager SOPS) code-quality gates (SonarQube) and CI/CD + GitOps (Jenkins ArgoCD/Helm). Benefits and Perks At impact.com we believe that when you’re happy and fulfilled you do your best work. That’s why we’ve built a benefits package that supports your well-being growth and work-life balance. Flexible Working Our Responsible PTO policy means you can take the time off you need to rest and recharge. We're committed to a positive work-life balance and provide a flexible environment that allows you to be happy and fulfilled in both your career and your personal life. Health and Wellness Your well-being is a priority. Our mental health and wellness benefit includes up to 12 fully covered therapy/coaching sessions per year with additional dependent coverage. We also offer a monthly gym reimbursement policy to support your physical health. A Stake in Our Growth We offer Restricted Stock Units (RSUs) as part of our total compensation giving you a stake in the company's growth with a 3-year vesting schedule pending Board approval. Investing in Your Growth We’re committed to your continuous learning. Take advantage of our free Coursera subscription and our PXA courses. Parental Support We offer a generous parental leave policy 26 weeks of fully paid leave for the primary caregiver and 13 weeks fully paid leave for the secondary caregiver. Technology Financial Support We provide a technology stipend to help you set up your home office and a monthly allowance to cover your internet expenses impact.com is proud to be an equal opportunity workplace. All employees and applicants for employment shall be given fair treatment and equal employment opportunity regardless of their race ethnicity or ancestry color or caste religion or belief age sex (including gender identity gender reassignment sexual orientation pregnancy/maternity) national origin weight neurodivergence disability marital and civil partnership status caregiving status veteran status genetic information political affiliation or other prohibited non-merit factors. #Li-Hybrid_CapeTown