All times are in the US Pacific timezone.
On Tuesday September 2, 2025 from 9:35am to 10:39am, Rippling experienced a major outage that left users unable to access the Rippling application. The outage impacted most users on the platform.
The outage was caused by two incompatible changes made in separate components of our system. Together, these changes triggered an infinite loop of database queries, overwhelming a critical database and dramatically increasing site latency. In an attempt to accelerate recovery, we enabled a database rate-limiting tool – meant to shed load from overloaded databases – which left the app unstable even after the original issue was resolved and prolonged the outage.
Rippling takes site reliability seriously and we apologize for this incident. Below, we will explain in detail how the issue arose, how it bypassed our safety procedures, and the future steps we are taking to prevent similar issues from occurring.
The scope of this incident spanned three separate subsystems:
Authentication
Rippling’s suite of products is built on a common set of components, including an authentication (“auth”) framework. Since the auth framework runs at the beginning of every API request, slowdowns in auth affect all products. The performance of the auth framework is critical for user experience.
Database ORM
An Object-Relational Mapper (“ORM”) is a tool that acts as a translator between application code and database queries, which often simplifies database access. Like many complex applications, Rippling uses an ORM library to manage queries to our databases, which we further customize to centralize database access patterns.
Feature flags
Rippling uses a feature flag system to safely deliver changes to production. The feature flag system allows us to selectively enable product features for specific users, and quickly rollout (or rollback) changes within seconds. The feature flag system automatically falls back to a safe default on errors.
On Aug 31, an Engineer made a change to our database ORM, adding an audit event for a specific data access pattern that we plan to change. This audit event was controlled by the feature flag system, introducing a new dependency not previously used in the ORM. This change was viewed as low-risk by code reviewers, in part because the feature flag system itself is a risk mitigation tool. The change was validated and deployed without issue.
On Sept 2, a second Engineer made a change to the feature flag system, designed to improve feature rollout targeting among users. This second change added a new query to the auth database (in a step called “build context”). The second Engineer was not aware of the previous change or that these changes would be incompatible. This change was validated by our automated tests, but upon further review we discovered an error in the corresponding test which made it ineffective.
When deployed, the new rollout targeting change invoked the previous ORM change. Because both components now referenced each other, the system entered an infinite loop and rapidly re-ran the new auth query due to the feature flag system’s automatic fallback.

We deploy all new versions of software at Rippling on a trial basis, known as our “canary”. The canary receives a small fraction of production traffic and automatically rolls back any new version of software with an elevated error count. This new version increased latency but decreased the number of requests, suppressing the error count and passing the error count test. Since our canary does not also test latency, the new version proceeded to a full deployment.
Once fully deployed, the new software version generated more queries than the auth database could handle. Because the auth database is used widely across Rippling, latency increased quickly across all products. This latency made the application effectively unusable.

Our Engineering team was able to quickly identify the issue and began to roll back the faulty deploy.
The rollback process shares configuration with our standard production deployment. Because this process is optimized for safety, deployments (including rollbacks) are incremental and can be slow. Attempting to accelerate recovery, we enabled a database rate-limiting tool to shed load from the auth database. However, this change increased error rates and further destabilized the system. Engineers identified this secondary issue and disabled the rate-limiting, allowing the app to fully recover.
Rippling has a policy to avoid deploying critical infrastructure changes during peak business hours (weekdays, 5am to 5pm), intended to prevent faulty deployments from negatively affecting customers. This issue bypassed that policy due to a misunderstanding around automated enforcement.
We pursue automation whenever possible – most parts of our deployment infrastructure are automatic. The second Engineer correctly recognized that their change affected critical infrastructure and believed our automated deployment process would schedule it accordingly. Unfortunately, the corresponding component was not marked as critical in our inventory and the change was deployed outside our policy window.
We have made immediate changes to our tools and policies:
Based on our learnings from this incident, we’re introducing the following changes to our infrastructure, process, and code:
This incident lasted nearly an hour and had a wide impact on Rippling customers. Ultimately, our automations failed to adequately protect our production deployment.
Again, we sincerely apologize for the disruption. We have begun the remediation steps above and will continue to address them urgently. As always, we will keep investing in our platform to further increase the reliability of our product.