“We’ll Clean It Up Later”: The Hidden Cost of the Snowflake Tables Nobody Remembers
Every data team has had this moment. Someone’s poking around the Snowflake catalog and stumbles on a table with a name like staging_backup_v2_final or temp_export_q3. Nobody recognizes it. Nobody’s sure who made it, what it was for, or whether anything downstream still quietly depends on it. So it stays. “We’ll clean it up later” becomes the answer, and later, reliably, never comes.
Multiply that moment by every table like it sitting in a real production Snowflake environment, and you start to see the actual shape of the problem. It’s not one table. It’s a slow accumulation of small, individually forgivable decisions to not deal with something right now, and it adds up to real money and real risk.
It’s not just a cost problem, even though it is one
The most obvious cost of an unused table is the storage bill, and that’s real. Storage isn’t free, and a table nobody’s touched in months is pure cost with zero value behind it. But if you only think about it as a line item, you’re missing most of the actual problem.
An old, forgotten table isn’t neutral just because nobody’s using it. It’s still sitting in your catalog, still showing up in searches, still a candidate for someone new on the team to stumble across and mistake for something current. A junior analyst building a new query doesn’t know that customer_data_v1 was superseded eighteen months ago; they just see a table with a plausible name and start querying it. Now you have a report built on stale or duplicated data, and nobody notices until the numbers don’t add up somewhere important.
That’s the quieter cost: every unused table is a small tax on everyone else’s trust in the catalog. The more of them there are, the harder it is to know, at a glance, which tables are the real, current source of truth.
Why “later” never actually comes
This isn’t a discipline problem. Data teams aren’t lazy about cleanup, they’re busy, and cleanup competes for time against everything that’s actually on fire today. Worse, cleaning up a table carries a specific kind of fear that new feature work doesn’t: what if it’s not actually dead? What if some dashboard, some scheduled job, some analyst’s personal script quietly depends on it, and dropping it breaks something nobody thought to check?
That fear is completely reasonable, and it’s exactly why “later” wins by default. Investigating whether a table is truly safe to remove takes real effort: checking pipeline definitions, searching query history, asking around. Multiply that effort by every candidate table in a real environment, and the honest, rational choice becomes leaving it alone and moving on to something with a clearer payoff.
What actually breaks the cycle
The fix isn’t a company-wide cleanup sprint once a year, those get scheduled, get deprioritized, and quietly disappear from the calendar. What breaks the cycle is removing the two things that make cleanup hard: the uncertainty about whether a table is really safe to drop, and the effort of investigating it yourself.
That uncertainty is exactly where most tools stop short, because they can only tell you a table hasn’t been queried recently, which isn’t the same as knowing it’s safe to remove. A table can go unqueried directly and still be a live step in a pipeline that feeds something important further downstream. Answering that properly means actually knowing how every table connects to everything else, which pipelines produce it, which ones read from it, across dbt, Airflow, Fivetran, and the rest of your stack. That’s Seemore’s own lineage, mapped across your whole environment, and it’s what lets a table only get flagged when it’s confirmed dead on both fronts: no pipeline connection anywhere, and no real query activity either. Not a guess based on one signal, an answer grounded in how your data actually flows.
That same care applies to the usage side of the check, not just the lineage side. A table can look “active” in query history simply because an automated bot or service account keeps pinging it, or because a specific person’s queries against it don’t actually reflect real, ongoing use. Seemore lets you exclude specific users, bots, and service accounts from counting as real usage, so a table like that still surfaces as the orphan it actually is, instead of hiding behind activity that was never meant to matter.
Once that uncertainty is gone, the only thing left is effort, and that’s solved by handing over the exact command to remove the table instead of leaving someone to write and double-check the SQL themselves, and by treating a “safe to delete” table the same way you’d treat any other piece of real work: assigned to someone, tracked to completion, not a report that gets glanced at once and forgotten.
Once cleanup stops being an investigation and starts being a checklist someone can actually work through, it stops competing with everything else on fire. It just becomes something that gets done.
What you get back
The dollar savings from freed storage are the easiest thing to point to, and they’re genuinely worth having. But the bigger win is a catalog people can actually trust: fewer tables that might be dead, might be live, nobody’s quite sure, and correspondingly less risk of someone building on the wrong foundation without realizing it. A cleaner environment isn’t just cheaper. It’s easier to reason about, easier to onboard new people into, and a little less likely to surprise you later.
Frequently asked questions
Why do unused tables pile up in the first place?
Not from carelessness, usually just from the normal churn of projects, migrations, and one-off exports. Each individual table felt harmless to leave behind at the time; the accumulation is what becomes the problem.
Is an unused table really costing us money if nobody's using it?
Yes. Storage isn’t free, so a table with zero activity is pure cost with no value behind it, even if it’s small on its own. The real expense usually comes from many small ones adding up.
Why is deleting a table so much scarier than it sounds?
Because the risk is asymmetric: leaving a useless table costs a little money, but deleting the wrong one can break a report or pipeline nobody remembered depended on it. That fear is exactly why cleanup keeps getting postponed.
Is data cleanliness really about more than storage cost?
Yes. An old, forgotten table can still get mistaken for something current, especially by someone new to the team, which can lead to decisions or reports built on stale or duplicated data without anyone noticing.
How do you make sure a table is actually safe to remove, not just quiet?
By checking two things together: whether it’s connected to any pipeline anywhere in your environment, and whether it’s had any real query activity. The pipeline check comes from Seemore’s own lineage, which maps how every table actually connects across dbt, Airflow, Fivetran, and the rest of your stack, not just whether someone happened to query it recently. A table only gets flagged when both checks come back clean.
What if a table only shows activity from a bot, service account, or a specific user I don't want counted?
You can exclude specific users, bots, or service accounts from counting as real usage. That way, a table kept artificially “active” by activity that doesn’t reflect genuine, ongoing use still gets correctly surfaced as an orphan.
What actually gets teams to follow through on cleanup, instead of just noting it and moving on?
Removing the investigation work. When a finding already tells you it’s safe, shows the exact command to run, and can be assigned and tracked like any other task, it competes on equal footing with everything else on the to-do list, instead of losing by default.