Abstract
Increasing demand for both greater parallelism and faster clocks dictate that future generation architectures will need to decentralize their resources and eliminate primitives that require single cycle global communication. A Raw microprocessor distributes all of its resources, including instruction streams, register files, memory ports, and ALUs, over a pipelined two-dimensional mesh interconnect, and exposes them fully to the compiler. Because communication in Raw machines is distributed, compiling for instruction-level parallelism (ILP) requires both spatial instruction partitioning as well as traditional temporal instruction scheduling. In addition, the compiler must explicitly manage all communication through the interconnect, including the global synchronization required at branch points. This paper describes RAWCC, the compiler we have developed for compiling general-purpose sequential programs to the distributed Raw architecture. We present performance results that demonstrate that although Raw machines provide no mechanisms for global communication the Raw compiler can schedule to achieve speedups that scale with the number of available functional units.
Cite
CITATION STYLE
Lee, W., Barua, R., Frank, M., Srikrishna, D., Babb, J., Sarkar, V., & Amarasinghe, S. (1998). Space-Time Scheduling of Instruction-Level Parallelismon a Raw Machine. Operating Systems Review (ACM), 32(5), 46–57. https://doi.org/10.1145/384265.291018
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.