Distributed Data and Computation · checkpoint worksheet

1 · Bottleneck

Workflow:
Data size / rows:
Elapsed time:
Plan or system clue:
Simplest plausible fix:

2 · Map / shuffle / reduce

Partition key:
Map output:
Shuffle key:
Reduce output:
Likely skew:
Why PostgreSQL is or is not already sufficient:

3 · Scale recommendation

Choose: one database / object-storage partitions / distributed compute
Evidence:
Acceptance threshold:
New failure mode introduced:
Trigger to reconsider: