Yes, you can hold p95 under 2 seconds at peak, 1,000+ QPS concurrency for your customer-facing dashboards.
A fintech unicorn did it with p95 under 2 seconds, without doubling the bill. Here is the problem as we are hearing it, what it costs, and the way out.
THE PROBLEM
Your query engine cannot hold p95 under 2 seconds at peak concurrency, so the dashboard you want to sell buckles at exactly the moment merchants are watching.
THE PROBLEM ANATOMY
Who it hits
- Data platform lead: owns the engine behind the merchant dashboards.
- Product owner: owns the SLA promised to merchants.
- Finance partner: owns the scaling bill.
The situation
A high-growth B2B fintech with 18,000+ merchants and petabytes of transaction data saw a revenue line sitting in its own warehouse: self-service analytics dashboards, sold to the merchants the data describes. At that scale a dashboard stops being an add-on. It has to hold strict latency SLAs on every query, stay elastic through bursts, and read near-real-time data.
The existing engine could not get there. In test runs, infrastructure cost ballooned as capacity scaled up, and performance still lagged. As load approached 1,000 QPS the legacy stack began to buckle, with p95 drifting past the 2-second threshold the product needed.
The specific challenges
- Peak loads near 1,000 QPS pushed p95 past the 2-second SLA
- Adding capacity raised the bill faster than it raised throughput
- Merchant insights had to come from fresh transactional data, not yesterday's batch
WHAT IT COSTS
PRACTICAL
- Test runs burned budget without reaching the SLA
- Engineers tuned a stack that could not hold concurrency and latency at once
- Every load spike meant crashes or emergency capacity
BUSINESS
- A revenue-generating analytics product stuck at pilot
- A slow, unreliable dashboard would sink the offering's value with merchants
EMOTIONAL
- The team could see the product, and the engine kept saying no
- Every capacity increase felt like paying twice for the same miss
THE WAY OUT
A lakehouse compute engine built for high QPS
The fintech team piloted e6data alongside their existing compute engines for the selected use case. The pilot offered three core propositions.
High concurrency support: the platform could scale to handle thousands of customers and their requirements without performance degradation or additional costs.
Autoscaling for peak workloads: an elastic execution layer that autoscales granularly to handle bursts of query traffic, then scales back down to save resources whenever possible.
Sub-second insights: the architecture enabled near-real-time querying on fresh transactional data, so merchants could get up-to-the-second insights.
The dashboards now ship as a premium merchant feature: a new revenue stream on data the company already had.
"With e6data in the mix, we finally hit p95 <2s on our customer-facing dashboards without doubling our bill. It's rare to see cost go down and performance jump that dramatically."
Book a demo on your own workloads
Reach out to book a demo, share challenges you're facing, and tell us how this fits into what you're currently working on or thinking about.
Prefer to self-serve? Problems we're solving →