← ALL PROBLEMS PERFORMANCE CLIFFS AND SCALE FAILURES

Yes, you can hold p95 under 2 seconds at peak, 1,000+ QPS concurrency for your customer-facing dashboards.

A fintech unicorn did it with p95 under 2 seconds, without doubling the bill. Here is the problem as we are hearing it, what it costs, and the way out.

ENVIRONMENT · AWS, S3, GLUE CATALOG

THE PROBLEM

Your query engine cannot hold p95 under 2 seconds at peak concurrency, so the dashboard you want to sell buckles at exactly the moment merchants are watching.

THE PROBLEM ANATOMY

Who it hits

  • Data platform lead: owns the engine behind the merchant dashboards.
  • Product owner: owns the SLA promised to merchants.
  • Finance partner: owns the scaling bill.

The situation

A high-growth B2B fintech with 18,000+ merchants and petabytes of transaction data saw a revenue line sitting in its own warehouse: self-service analytics dashboards, sold to the merchants the data describes. At that scale a dashboard stops being an add-on. It has to hold strict latency SLAs on every query, stay elastic through bursts, and read near-real-time data.

The existing engine could not get there. In test runs, infrastructure cost ballooned as capacity scaled up, and performance still lagged. As load approached 1,000 QPS the legacy stack began to buckle, with p95 drifting past the 2-second threshold the product needed.

The specific challenges

  • Peak loads near 1,000 QPS pushed p95 past the 2-second SLA
  • Adding capacity raised the bill faster than it raised throughput
  • Merchant insights had to come from fresh transactional data, not yesterday's batch

WHAT IT COSTS

PRACTICAL

  • Test runs burned budget without reaching the SLA
  • Engineers tuned a stack that could not hold concurrency and latency at once
  • Every load spike meant crashes or emergency capacity

BUSINESS

  • A revenue-generating analytics product stuck at pilot
  • A slow, unreliable dashboard would sink the offering's value with merchants

EMOTIONAL

  • The team could see the product, and the engine kept saying no
  • Every capacity increase felt like paying twice for the same miss

THE WAY OUT

A lakehouse compute engine built for high QPS

The fintech team piloted e6data alongside their existing compute engines for the selected use case. The pilot offered three core propositions.

01

High concurrency support: the platform could scale to handle thousands of customers and their requirements without performance degradation or additional costs.

02

Autoscaling for peak workloads: an elastic execution layer that autoscales granularly to handle bursts of query traffic, then scales back down to save resources whenever possible.

03

Sub-second insights: the architecture enabled near-real-time querying on fresh transactional data, so merchants could get up-to-the-second insights.

1,000+
QPS sustained under load
<2s
p95 latency, complex joins included
50%+
lower TCO than the previous engine
12x
faster query completion (mix of OLAP and near real-time workloads)

The dashboards now ship as a premium merchant feature: a new revenue stream on data the company already had.

"With e6data in the mix, we finally hit p95 <2s on our customer-facing dashboards without doubling our bill. It's rare to see cost go down and performance jump that dramatically."

STAFF DATA ENGINEER · FINTECH UNICORN

Book a demo on your own workloads

Reach out to book a demo, share challenges you're facing, and tell us how this fits into what you're currently working on or thinking about.

Prefer to self-serve? Problems we're solving

An actual person replies. By submitting, you acknowledge your personal information will be processed in accordance with our Privacy Policy.