< blog
8 min read

The Silent Snowflake Cost Leak: Warehouses That Never Got Resized Back

Quick answer: We see this over and over: the costliest anomaly on a Snowflake bill is almost never a bad query. It’s a good decision nobody walked back — a warehouse scaled up under pressure, then just… left. It’s less a Snowflake problem than a human one: it’s always easier to make an exception than to remember you made it. The fix isn’t a better reminder. It’s building the ending into the decision itself.


Here’s the story. It’s always some version of the same story.

It’s a Thursday night, or a Monday before a board meeting, or the middle of a launch nobody planned properly. A backfill’s crawling. A dashboard’s timing out. A report is due in an hour and it is very much not going to make it at the current pace. So someone does the only sensible thing: they bump the warehouse up. Small to Medium, Medium to Large – whatever it takes. Ten minutes later the job finishes, the meeting goes fine, the fire’s out. And that’s it. Nobody thinks about that warehouse again.

We see this over and over. Not as a one-off horror story a customer tells us once – as a pattern that shows up again, and again, and again, every single time we sit down and actually look at where the money went. A warehouse jumps in size on some unremarkable Tuesday, holds there through the weekend, and is still holding a month later — quietly billing away in the background like nothing ever happened, because as far as the system is concerned, nothing did.

The upgrade is a sprint. The downgrade is homework nobody assigned.

Here’s what we keep noticing: scaling up is the easiest decision in the world. One click, real urgency, instant payoff – the job finishes, the dashboard loads, the exec stops refreshing the page. Scaling back down has none of that. No deadline. No applause. No one on the other end waiting for it. It’s homework with an infinite due date, competing for headspace against literally everything else on someone’s plate. And homework with no due date loses.

And this is the part that should needle anyone who’s spent real time in ops: it’s not really about warehouses. We’ve watched the exact same shape show up everywhere – the feature flag nobody flipped back off, the firewall rule opened “just for the demo” two quarters ago, the elevated access granted for one debugging session that, somehow, is still very much active. Wherever a system lets you make a fast, technically-reversible change, the reversal ends up with no owner and no clock – so it just doesn’t happen. Snowflake just happens to hand you the receipt in dollars, itemized, every single month, whether or not anyone’s reading it.

Alerts don’t stop the bleeding. They just tell you it happened.

The instinct, every time, is to watch harder. A cost anomaly alert. A tagging convention. A monthly Slack nudge to “please right-size your warehouses, team.” We’ve seen all of it tried. All of it well-meaning. None of it touching the actual cause.

By the time an alert fires, the money’s already spent. A monthly review catches whatever’s still standing when someone finally looks, and misses everything that happened in between. And a policy that says “please remember” is really just asking people to be disciplined in the exact moment they’ve already proven, over and over, that they won’t be – mid-deadline, mid-fire, mid-everything-else. You cannot out-process a decision that was never going to get a second thought to begin with.

Stop asking people to remember. Make forgetting impossible.

So here’s the idea that actually seems to hold up, the one that keeps surfacing once you’ve watched this happen enough times: flip the default. Don’t build a change that lives forever unless someone remembers to kill it. Build one that’s temporary from the second it’s created, and make permanence the thing that requires a deliberate choice – not the other way around. Put the expiration date on the decision itself, right when it’s made, instead of hoping someone circles back later with a clear head.

It’s not even a Snowflake idea, really. It’s the same logic behind time-boxed access grants and auto-expiring flags in any system built by people who’ve been burned before. Applied to a warehouse, it means the system asks “for how long?” the instant someone reaches for more power – and then it actually holds them to the answer. No hoping. No ticket quietly dying in a backlog. No trust required, because none is needed.

This is exactly the itch that led us to build Live Tuning

Activate Live Tuning dialog: choose a duration, warehouse size and warehouse type

Once you’ve watched this play out enough times, you stop wanting to warn people about it and start wanting to make it structurally impossible. That’s Live Tuning. Every manual change to a warehouse – size or type – comes with a built-in expiration, anywhere from 5 minutes to 24 hours, chosen in the same breath as the change itself. When the clock runs out, the warehouse reverts on its own. Back to its SmartPulse schedule, if it has one. Back to its saved configuration, if it doesn’t. Nobody has to remember anything, because there’s nothing left to remember.

Nothing here happens in the dark

Real-time warehouse view in Seemore with an active Live Tuning session

Here’s the other half of it: none of this runs on faith. The whole point is that you’re never left wondering what’s actually going on with a warehouse mid-session. Real Time Monitoring shows an active Live Tuning session live – the countdown, the size, the configuration, exactly like watching a timer you set yourself. And every single thing that happens to it – who started it, what they changed, when it ended, whether it expired on its own or got stopped early – lands in the Event Log, attributed to the person who made the call. No mystery resize, no “wait, who did this and why.” If a warehouse changed shape, there’s a name and a timestamp next to it.


FAQ

What actually causes most unexplained Snowflake cost spikes?
Almost never a bad query, almost never real growth. What we see, over and over, is a manual warehouse change made for a genuinely good reason – and simply left running long after that reason was gone.

Why don’t cost alerts fix this?
Because they show up after the story’s already over. By the time an alert fires, the cost is sunk, and the moment someone could’ve reverted the change is long past.

What’s the real fix, structurally?
Build the expiration into the decision itself instead of trusting anyone to remember it later – the same logic behind time-boxed access grants or auto-expiring feature flags, applied to compute.

How does Seemore Data’s Live Tuning feature put that into practice?
Any manual override – size or type – gets a duration attached the moment it’s made, from 5 minutes up to 24 hours. When time’s up, it reverts automatically: to SmartPulse if one’s running, or to the saved configuration if not.

Does this replace something like SmartPulse?
No – different jobs, and we see both show up in the same account all the time. SmartPulse handles the patterns you can predict. Live Tuning handles the ones you can’t, without leaving a loose end behind.

Is this really just a Snowflake thing?
Not even close – it’s just where we happen to see it. It’s what happens anywhere a system lets you make a fast, “reversible” change without building in the reversal. Snowflake just hands you the invoice.

Cool, now
what can you DO with this?

data ROI