Technical question · Pipelines
Explain idempotency to a new analyst on your team, using a job that runs every night as the example.
What they’re really asking
The interviewer wants to know that you understand one of the ideas that keeps pipelines trustworthy, and that you can explain it without jargon. They’re listening for a plain definition, a nightly job that shows why it matters, what goes wrong when a job isn’t idempotent, and how you’d build one that is. Pitching it at a new analyst rather than a fellow engineer is part of the test.
How to structure your answer
- Define it plainlyRunning the job twice with the same input leaves the data exactly as running it once would.
- Say why it mattersJobs get rerun all the time: after a failure, an automatic retry, a bug fix or a backfill.
- Show what goes wrongA job that simply adds rows doubles the day’s figures when it’s run again.
- Explain how to build itReplace the day’s data in one step, or merge on a key, instead of adding rows blindly.
For more on answering this kind of question, including the mistakes to avoid, read our guide to technical questions.
An example outline
This is an example outline, not a script. Use it to see the shape of a strong answer, then build yours from your own experience, in your own words.
- Definition
- Idempotent means you can run the job again and get the same result, not a second copy of it.
- The job
- Every night at 2am, a job loads yesterday’s sales from the tills into the sales table.
- What goes wrong
- One night it fails halfway and someone reruns it. If it only adds rows, the half that loaded the first time is now in the table twice, and Monday’s sales look higher than they were.
- The fix
- The job loads a given day, not “whatever is new”. It deletes that day’s rows and reloads them in one transaction, or overwrites that day’s partition. Run it once or five times: same table.
- Why it matters to you
- A rerun becomes a safe, boring fix, and the numbers you report don’t change just because a job ran twice.
Related questions
- A supplier sends you a price file every day. Sometimes they send the same file twice, and sometimes a corrected version of an earlier day’s file. How would you design the load?Pipelines · 5 minutes to answer
- When is a full reload better than an incremental load, and when does it stop being affordable?Pipelines · 2 minutes to answer
- You need to backfill three years of history into a table that already feeds live dashboards. How would you do it safely?Pipelines · 5 minutes to answer
- When would you process data in batches, and when would you stream it? Give an example of each.Pipelines · 2 minutes to answer
Read the guide
- How to answer a data pipeline design question as a junior data engineerInterview guide · 7 min read
- How to talk through a data project in an interviewInterview guide · 7 min read
- How to explain your SQL out loud in an interviewInterview guide · 7 min read
Last reviewed 28 September 2026 by the Deeplink Interview team. How we write and check our questions.