Resolved
All data has been recovered and the system health is back to normal
Monitoring
We are still getting rate-limited by our upstream storage provider but we are draining retries faster than before because of some optimizations we have made on the system. We are working on accelerating them as much as possible but we still have some retries in the queue.
No data has been lost, all affected inserts are being safely retried.
Monitoring
The bulk of the earlier retries has now been recovered, and no new ingestion errors are being introduced on the customer side. Recovery is progressing more slowly than we'd like because we are currently being rate-limited by our upstream storage provider, which is limiting how quickly we can drain the remaining retries.
No data has been lost, all affected inserts are being safely retried. We'll provide the next update as soon as we see meaningful progress on either front.
Monitoring
We're continuing to work through a backlog of delayed ingestion on our AWS US East 1 shared cluster following yesterday's degradation. Our coordination layer was resized last night and the bulk of the impact has cleared. Recovery is taking longer than expected, so affected workspaces may continue to see ingestion lag while we work through the remaining queue.
Monitoring
We keep ingesting the delayed real time data
Monitoring
We've identify the root cause and we are recovering the retries from the HFI ingestion at the moment. The platform should be stable now
Identified
We've already identified the root cause and we're working through the fix. Real-time ingestion is now fully available, we're working on ingesting past data that hasn't been processed yet.
Identified
We're still working on stabilising the region. We're struggling with Zookeeper, in charge of coordination between ClickHouse replicas, which is why write are a lot more affected than reads.
We've just tweaked a couple parameters that hope will aid the situation.
Investigating
We keep working on identifying the root cause. There's impact in both reads and writes, and in all clusters in the region as the incident is affecting a shared component.
We are storing the data that gets to the Events API to retry the inserts later.
Investigating
There’s a performance degradation in the public cluster on AWS us-east.