Multiprocessors and vector machines, the only successful parallel architectures, have coarse-grained parallelism that is hard for compilers to take advantage of. We've developed a new fine-grained parallel architecture and a compiler that together offer order-of-magnitude speedups for ordinary scientific code.
No takes yet. Share an insight, caveat, or question.
Fisher et al. (1984) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: