Every Sunday, I take 15 minutes to look back at the week — not just what I built, but how I thought. This habit has quietly become one of the most valuable things I do for my career.
Every week in data engineering teaches you something new — often the hard way. This post collects the hardest lessons from two of those weeks: inheriting an over-engineered Airflow DAG, refactoring a dbt model that had grown way too clever for its own good, and explaining a data discrepancy to a frustrated stakeholder. The biggest wins came from doing less, not more.
If you’re in the data space — whether you’re just starting out or have been building pipelines for years — I think you’ll recognise at least one of these.
1. Simpler Pipelines Beat Clever Ones (Almost Always)
I inherited an Airflow DAG that had 14 tasks, custom sensors, dynamic task mapping, and enough conditional logic to make your head spin. It was impressive — but it was also breaking constantly and nobody could debug it in under an hour.
We replaced it with a dbt model + a single cron job. Result: 80% less code, same output, and any junior engineer on the team can now understand and maintain it.
There’s a temptation in data engineering to show off. To write the cleverly nested SQL that handles 12 edge cases in one subquery. To build the Spark job that processes everything in a single stage. But cleverness has a cost: maintainability. When I revisited a pipeline I wrote six months ago, I spent 45 minutes figuring out what I was trying to do. The fix? Break it into smaller dbt models, add a comment explaining why, and drop the clever tricks. A junior engineer can now understand it in under five minutes.
Complexity is not sophistication. Simplicity is a feature, not a shortcut. Rule of thumb: if your pipeline would confuse a smart colleague who hasn’t seen it before — or needs a presentation to explain it — it’s too complex.
2. Query Execution Plans Are Underrated
I started spending 30 minutes each morning reviewing EXPLAIN ANALYZE output on our slowest queries. Within three days, I found two silent killers: a full table scan on a 200M-row table and a nested loop join picking the wrong strategy due to stale statistics.
EXPLAIN ANALYZE
SELECT *
FROM orders o
JOIN customers c ON o.customer_id = c.id
WHERE o.created_at > NOW() - INTERVAL '7 days';
Takeaway: Reading execution plans feels slow. Not reading them is slower.
3. Always Ask “Why Do We Need This?”
Before writing a single line of code, ask your stakeholder: What decision will this data enable? Who will use it? How often?
A stakeholder came to me with a “quick” request: connect 3 new data sources. Old me would’ve said yes. This time I asked those questions. The answers were vague. The request got deprioritized.
Another request came in for a new aggregation table. After two minutes of questions, it turned out the stakeholder just wanted a number already available in an existing dashboard. Two hours of engineering work avoided with a two-minute conversation.
Every new data source is a long-term maintenance commitment. Be selective. A lean data platform that reliably serves 10 use cases is worth more than a sprawling one that partially serves 50.
4. Documentation Debt Is Real — Document at the Point of Understanding
I came back to a Python utility script I wrote 6 weeks ago. No comments. No README. No docstrings. I spent 45 minutes reverse-engineering what I had written.
def normalize_event_timestamps(df: pd.DataFrame, tz: str = "UTC") -> pd.DataFrame:
"""
Convert all timestamp columns to a unified timezone.
Args:
df: Input DataFrame with raw event data
tz: Target timezone string (default: 'UTC')
Returns:
DataFrame with normalized timestamp columns
"""
# implementation here
A docstring + type hints. Takes 2 minutes. Saves 45 minutes later.
We all know documentation matters. We all leave it for later. Later never comes. So I tried documenting at the moment I understood something. It added 10 minutes to my day but saved a 30-minute “how does this work again?” on Friday afternoon.
Practical tip: In dbt, treat the description: field in schema.yml as mandatory. A single sentence explaining why a model exists is more valuable than a paragraph describing what SQL it runs.
5. Data Quality Issues Are Communication Issues
When a number is wrong, the instinct is to dive into SQL. But more often, the root cause isn’t technical — it’s a misunderstanding between engineering and the business about what a metric actually means. One “data quality issue” turned out to be a disagreement about whether returns should be excluded from revenue before or after a certain date. Not a bug. A definition problem.
The fix: Before you debug, align on the definition. A shared data dictionary isn’t a luxury — it’s infrastructure.
6. Knowing When to Stop Is a Skill
I spent three hours optimising a Spark job and got it 15% faster. Was that worth three hours? Probably not — the job ran once a day and the business didn’t care. Premature optimisation in data engineering is just as dangerous as in software engineering. Know your SLAs. Save your energy for the jobs that actually matter.
The Mindset Shift That Ties It Together
Stop asking “how do I build this?” Start asking “should I build this at all?”
Most data problems are not engineering problems. They’re clarity problems. The best data engineers push back — not to be difficult, but to make sure the work they do actually matters.
Wrapping Up
The fundamentals don’t change: build things simply, understand the problem before you build, communicate clearly, and know when to stop.
If you’re a data engineer, spend 15 minutes every Sunday asking: What worked and why? What didn’t work and what would I do differently? What’s one thing I’ll carry into next week?
Small habit. Big compounding returns. If this week was tough, you’re in good company. Keep going.
What was your biggest lesson this week? I’d love to hear it in the comments.
— Pushpjeet Cholkar, Data Engineer
Newsletter
Enjoyed this post?
Get new posts, AI/ML tutorials, AWS batch dates and Oracle tips by email. No spam, unsubscribe anytime.