Distributed Data and Computation · checkpoint worksheet
1 · Bottleneck
Workflow:
Data size / rows:
Elapsed time:
Plan or system clue:
Simplest plausible fix:
2 · Map / shuffle / reduce
Partition key:
Map output:
Shuffle key:
Reduce output:
Likely skew:
Why PostgreSQL is or is not already sufficient:
3 · Scale recommendation
Choose: one database / object-storage partitions / distributed
compute
Evidence:
Acceptance threshold:
New failure mode introduced:
Trigger to reconsider: