| Author | Title | Year | Journal/Proceedings | Reftype | DOI/URL |
|---|---|---|---|---|---|
| DeVuyst, M., Venkat, A. and Tullsen, D.M. | Execution migration in a heterogeneous-ISA chip multiprocessor | 2012 | SIGARCH Computer Architecture News Vol. 40(1), pp. 261–272 |
article | DOI URL |
| Abstract: Prior research has shown that single-ISA heterogeneous chip multiprocessors have the potential for greater performance and energy efficiency than homogeneous CMPs. However, restricting the cores to a single ISA removes an important opportunity for greater heterogeneity. To take full advantage of a heterogeneous-ISA CMP, however, we must be able to migrate execution among heterogeneous cores in order to adapt to program phase changes and changing external conditions (e.g., system power state).This paper explores migration on heterogeneous-ISA CMPs. This is non-trivial because program state is kept in an architecture-specific form; therefore, state transformation is necessary for migration. To keep migration cost low, the amount of state that requires transformation must be minimized. This work identifies large portions of program state whose form is not critical for performance; the compiler is modified to produce programs that keep most of their state in an architecture-neutral form so that only a small number of data items must be repositioned and no pointers need to be changed. The result is low migration cost with minimal sacrifice of non-migration performance.Additionally, this work leverages binary translation to enable instantaneous migration. When migration is requested, the program is immediately migrated to a different core where binary translation runs for a short time until a function call is reached, at which point program state is transformed and execution continues natively on the new core.This system can tolerate migrations as often as every 100 ms and still retain 95% of the performance of a system that does not do, or support, migration. | |||||
| Comment: This paper addresses the challenge of migrating program execution across cores with different instruction set architectures (ISAs) in a single chip multiprocessor (CMP). The core problem is that program state, including registers and memory, is architecture-specific, making migration costly due to the need for state transformation. The authors propose compiler and runtime techniques to minimize this transformation cost by ensuring most program state remains architecture-neutral, thus reducing the need for data repositioning or pointer adjustments. Key contributions include modifying the compiler to maintain consistent memory layouts (e.g., identical function ordering, padding stack frames) and leveraging binary translation for efficient migration. The paper also introduces methods to handle pointers and system calls uniformly across ISAs. Experimental results using SPEC2000 benchmarks on an ARM-MIPS CMP show migration costs are significantly reduced, with performance overheads under 5%. New concepts include architecture-neutral state preservation and equivalence points for migration. The work demonstrates that frequent, low-cost migration is feasible in heterogeneous-ISA CMPs, enabling dynamic adaptation to program phases or system conditions. | |||||
BibTeX:
@article{DeVuyst2012,
author = {DeVuyst, Matthew and Venkat, Ashish and Tullsen, Dean M.},
title = {Execution migration in a heterogeneous-ISA chip multiprocessor},
journal = {SIGARCH Computer Architecture News},
publisher = {ACM},
year = {2012},
volume = {40},
number = {1},
pages = {261–272},
url = {/docs/DeVuyst2012.pdf},
doi = {https://doi.org/10.1145/2189750.2151004}
}
|
|||||
| Mittal, S. | A Survey of Techniques for Architecting and Managing Asymmetric Multicore Processors | 2016 | ACM Computing Surveys (CSUR) Vol. 48(3) |
article | DOI URL |
| Abstract: To meet the needs of a diverse range of workloads, asymmetric multicore processors (AMPs) have been proposed, which feature cores of different microarchitecture or ISAs. However, given the diversity inherent in their design and application scenarios, several challenges need to be addressed to effectively architect AMPs and leverage their potential in optimizing both sequential and parallel performance. Several recent techniques address these challenges. In this article, we present a survey of architectural and system-level techniques proposed for designing and managing AMPs. By classifying the techniques on several key characteristics, we underscore their similarities and differences. We clarify the terminology used in this research field and identify challenges that are worthy of future investigation. We hope that more than just synthesizing the existing work on AMPs, the contribution of this survey will be to spark novel ideas for architecting future AMPs that can make a definite impact on the landscape of next-generation computing systems. | |||||
| Comment: This paper is a comprehensive survey of research on Asymmetric Multicore Processors (AMPs) which combine cores of different sizes and performance characteristics. It addresses the problem of optimizing energy efficiency and performance for diverse workloads in modern computing systems. The main contributions include clarifying terminology, analyzing performance potential, and reviewing optimization techniques for both static and reconfigurable AMPs. The paper distinguishes between static AMPs with fixed core configurations and reconfigurable AMPs that can adapt their core mix at runtime. New concepts introduced include heterogeneous-ISA AMPs, federated cores combining simple cores to approximate out-of-order execution, and execution objects for managing reconfigurable systems. It discusses techniques such as DVFS, fuzzy logic scheduling, CPI stack analysis, and cross-ISA migration for performance optimization. The survey categorizes works based on optimization objectives including energy efficiency, performance, fairness, reliability, and domain-specific techniques. Results synthesized from numerous research works show AMPs can significantly improve energy efficiency and throughput while facing challenges in scheduling and reconfiguration overhead. The paper provides a taxonomy of approaches across 45:34 categories including energy management, fairness, performance predictability, and bottleneck acceleration. It concludes with a future outlook discussing challenges like fault tolerance due to process variation and the need for extreme power management in exascale computing systems. | |||||
BibTeX:
@article{Mittal2016,
author = {Mittal, Sparsh},
title = {A Survey of Techniques for Architecting and Managing Asymmetric Multicore Processors},
journal = {ACM Computing Surveys (CSUR)},
publisher = {ACM},
year = {2016},
volume = {48},
number = {3},
url = {/docs/Mittal2016.pdf},
doi = {https://doi.org/10.1145/2856125}
}
|
|||||
| Xing, T., Xiong, C., Wei, T., Sanchez, A., Ravindran, B., Balkind, J. and Barbalace, A. | Stramash: A Fused-Kernel Operating System For Cache-Coherent, Heterogeneous-ISA Platforms | 2025 | Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (APLOS), pp. 1172–1188 | inproceedings | DOI URL |
| Abstract: We live in the world of heterogeneous computing. With specialised elements reaching all aspects of our computer systems and their prevalence only growing, we must act to rein in their inherent complexity. One area that has seen significantly less investment in terms of development is heterogeneous-ISA systems, specifically because of complexity. To date, heterogeneous-ISA processors have required significant software overheads, workarounds, and coordination layers, making the development of more advanced software hard, and motivating little further development of more advanced hardware. In this paper, we take a fused approach to heterogeneity, and introduce a new operating system (OS) design, the fused-kernel OS, which goes beyond the multiple-kernel OS design, exploiting cache-coherent shared memory among heterogeneous-ISA CPUs as a first principle -- introducing a set of new OS kernel mechanisms. We built a prototype fused-kernel OS, Stramash-Linux, to demonstrate the applicability of our design to monolithic OS kernels. We profile Stramash OS components on real hardware but tested them on an architectural simulator -- Stramash-QEMU, which we design and build. Our evaluation begins by validating the accuracy of our simulator, achieving an average of less than 4% errors. We then perform a direct comparison between our fused-kernel OS and state-of-the-art multiple-kernel OS designs. Results demonstrate speedups of up to 2.1texttimes on NPB benchmarks. Further, we provide an in-depth analysis of the differences and trade-offs between fused-kernel and multiple-kernel OS designs. | |||||
| Comment: Stramash presents a fused-kernel OS design for cache-coherent heterogeneous-ISA platforms addressing fundamental hardware-software mismatches in integrating different ISAs. The core contribution is the fused-kernel concept where multiple kernel instances coordinate via shared memory under a shared-mostly principle rather than shared-nothing. This enables direct access to kernel data structures eliminating serialization, deserialization, and message passing overheads between kernels. New concepts include a single kernel-level virtual address space across heterogeneous-ISA cores and minimal secure kernel data sharing using hardware enforcement mechanisms. The implementation provides Stramash-Linux based on the Linux kernel with a cross-ISA page table lock for synchronization and uses the Popcorn-Linux Compiler Toolchain for application compilation and migration. Stramash-QEMU simulator is introduced featuring instruction count-based timing and cache modeling plugins for performance evaluation. Results demonstrate hardware cache coherence at cacheline granularity outperforms software DSM by over 300x for single cacheline access. Redis-server experiments show 4-12x speedup with Stramash-Linux compared to TCP/IP messaging while NPB benchmarks exhibit varying performance outcomes. The artifact includes source code, modified kernels, simulator, and helper scripts publicly available with MIT license. | |||||
BibTeX:
@inproceedings{Xing2025,
author = {Xing, Tong and Xiong, Cong and Wei, Tianrui and Sanchez, April and Ravindran, Binoy and Balkind, Jonathan and Barbalace, Antonio},
title = {Stramash: A Fused-Kernel Operating System For Cache-Coherent, Heterogeneous-ISA Platforms},
booktitle = {Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (APLOS)},
publisher = {ACM},
year = {2025},
pages = {1172–1188},
url = {/docs/Xing2025.pdf},
doi = {https://doi.org/10.1145/3676641.3716275}
}
|
|||||
| Cho, S., Chen, H., Madaminov, S., Ferdman, M. and Milder, P. | Flick: Fast and Lightweight ISA-Crossing Call for Heterogeneous-ISA Environments | 2020 | Proceedings of the 47th ACM/IEEE International Symposium on Computer Architecture (ISCA), pp. 187–198 | inproceedings | DOI URL |
| Abstract: Heterogeneous-ISA multi-core systems have performance and power consumption benefits. Today, numerous system components, such as NVRAMs and Smart NICs, already have built-in processor cores with ISAs different from that of the host CPUs, making many modern systems heterogeneous-ISA multi-core systems. Unfortunately, programming and using such systems efficiently is difficult and requires extensive support from the host operating systems. Existing programming solutions are complex, require dramatic changes to the systems, and often incur significant performance overheads. To address this challenge, we propose Flick: Fast and Lightweight ISA-Crossing Call, for migrating threads in heterogeneous-ISA multi-core systems. By leveraging hardware virtual memory support and standard operating system mechanisms, a software thread can transparently migrate between cores with different ISAs. We prototype a heterogeneous-ISA multi-core system using FPGAs with off-the-shelf hardware and software to evaluate Flick. Experiments with microbenchmarks and a BFS application show that Flick requires only minor changes to the existing OS and software, and incurs only 18ps round trip overhead for migrating a thread through PCIe, which is at least 23x faster than prior work. | |||||
| Comment: Flick addresses the difficulty of programming heterogeneous-ISA multi-core systems, in particular systems where devices such as Smart NICs, NVMe storage, or FPGAs embed general-purpose processors with an ISA different from the host CPU (the paper calls these NxPs, or Near-x-Processors). Existing approaches either force developers into the offload-engine style with manual data movement and separate binaries, or rely on general thread-migration frameworks whose overhead is reported to be hundreds of microseconds to milliseconds, which discourages fine-grained migration. The authors propose Flick, a mechanism that lets a thread migrate between cores of different ISAs at function-call boundaries, using hardware virtual memory support and standard OS mechanisms rather than binary translation or stack transformation. Key elements are a unified virtual and physical address space shared by host and NxP (the NxP maps host memory directly and exposes its local memory to the host through a PCIe BAR), page-fault-triggered migration when a thread attempts to execute code compiled for the other ISA, and a multi-ISA "fat" executable produced by a modified linker and kernel module loader with per-ISA .text sections aligned to 4KB pages. Migration handlers on both sides are reentrant, which the authors state allows nested and recursive bidirectional calls across ISAs. Required system changes are described as modest: a custom linker script and relocation support for both x86-64 and RISC-V, and a small Linux scheduler modification adding a "migration" flag in task_struct so the descriptor DMA is triggered only after the thread has been suspended, avoiding a race condition. The prototype is an Intel Xeon host connected over PCIe to an FPGA implementing an in-order scalar RV64I core with instruction/data TLBs and an MMU realized as a small programmable micro-controller that walks the host x86 page tables, which also permits configurable "holes" bypassing translation for local scratchpad access. Reported results give a round-trip migration cost of 18 microseconds through PCIe, which the authors claim is at least 23 times faster than prior work. In a breadth-first search application on graph datasets stored in NxP-side DRAM, with a migration back to the host for every newly discovered vertex, Flick is slower than a host-side PCIe traversal baseline on the smallest graph (high vertex-to-edge ratio) but 9% to 19% faster on the two larger graphs. I note that the excerpt provided is partial, so some evaluation details, the microbenchmark numbers, and the identity of the baseline prior work used for the 23x comparison could not be verified from the text shown. | |||||
BibTeX:
@inproceedings{Cho2020,
author = {Cho, Shenghsun and Chen, Han and Madaminov, Sergey and Ferdman, Michael and Milder, Peter},
title = {Flick: Fast and Lightweight ISA-Crossing Call for Heterogeneous-ISA Environments},
booktitle = {Proceedings of the 47th ACM/IEEE International Symposium on Computer Architecture (ISCA)},
publisher = {ACM/IEEE},
year = {2020},
pages = {187–198},
url = {/docs/Cho2020.pdf},
doi = {https://doi.org/10.1109/ISCA45697.2020.00026}
}
|
|||||
| Mavrogeorgis, N., Vasiladiotis, C., Mu, P., Khordadi, A., Franke, B. and Barbalace, A. | UNIFICO: Thread Migration in Heterogeneous-ISA CPUs without State Transformation | 2024 | Proceedings of the 33rd ACM SIGPLAN International Conference on Compiler Construction, pp. 86–99 | inproceedings | DOI URL |
| Abstract: Heterogeneous-ISA processor designs have attracted considerable research interest. However, unlike their homogeneous-ISA counterparts, explicit software support for bridging ISA heterogeneity is required. The lack of a compilation toolchain ready to support heterogeneous-ISA targets has been a major factor hindering research in this exciting emerging area. For any such compiler “getting right” the mechanics involved in state transformation upon migration and doing this efficiently is of critical importance. In particular, any runtime conversion of the current program stack from one architecture to another would be prohibitively expensive. In this paper, we design and develop Unifico, a new multi-ISA compiler that generates binaries that maintain the same stack layout during their execution on either architecture. Unifico avoids the need for runtime stack transformation, thus eliminating overheads associated with ISA migration. Additional responsibilities of the Unifico compiler backend include maintenance of a uniform ABI and virtual address space across ISAs. Unifico is implemented using the LLVM compiler infrastructure, and we are currently targeting the x86-64 and ARMv8 ISAs. We have evaluated Unifico across a range of compute-intensive NAS benchmarks and show its minimal impact on overall execution time, where less than 6% overhead is introduced on average. When compared against the state-of-the-art Popcorn compiler, Unifico reduces binary size overhead from ∼200% to ∼10%, whilst eliminating the stack transformation overhead during ISA migration. | |||||
| Comment: The paper addresses the programmability challenge of heterogeneous-ISA platforms where traditional single-ISA compilation fails and thread migration across ISAs is needed but hindered by ABI and architectural differences. Emerging hardware such as CXL enables coherent shared memory across diverse ISAs like x86 and ARM, making thread migration more feasible than traditional offloading. The authors propose UNIFICO, a system enabling thread migration in heterogeneous-ISA CPUs without requiring state transformation of stack frames. UNIFICO is implemented using the LLVM compiler infrastructure with modifications to instruction selection, register allocation, and optimization passes. The system was evaluated on NPB and SPEC CPU2017 benchmarks, demonstrating average no more than 10% binary size overhead and no more than 6% execution time overhead without migration. A key innovation is the introduction of aligned stack layouts across architectures to avoid state transformation during migration. The paper also details compiler modifications to bridge differences in register allocation decisions and optimization behaviors between ISAs. UNIFICO is released as open-source software to facilitate adoption. The work demonstrates that heterogeneous-ISA thread migration is practical with minimal performance and size overheads. | |||||
BibTeX:
@inproceedings{Mavrogeorgis2024,
author = {Mavrogeorgis, Nikolaos and Vasiladiotis, Christos and Mu, Pei and Khordadi, Amir and Franke, Björn and Barbalace, Antonio},
title = {UNIFICO: Thread Migration in Heterogeneous-ISA CPUs without State Transformation},
booktitle = {Proceedings of the 33rd ACM SIGPLAN International Conference on Compiler Construction},
publisher = {ACM},
year = {2024},
pages = {86–99},
url = {/docs/Mavrogeorgis2024.pdf},
doi = {https://doi.org/10.1145/3640537.3641565}
}
|
|||||
| Venkat, A. and Tullsen, D.M. | Harnessing ISA diversity: design of a heterogeneous-ISA chip multiprocessor | 2014 | SIGARCH Computer Architecture News Vol. 42(3), pp. 121–132 |
article | DOI URL |
| Abstract: Heterogeneous multicore architectures have the potential for high performance and energy efficiency. These architectures may be composed of small power-efficient cores, large high-performance cores, and/or specialized cores that accelerate the performance of a particular class of computation. Architects have explored multiple dimensions of heterogeneity, both in terms of micro-architecture and specialization. While early work constrained the cores to share a single ISA, this work shows that allowing heterogeneous ISAs further extends the effectiveness of such architecturesThis work exploits the diversity offered by three modern ISAs: Thumb, x86-64, and Alpha. This architecture has the potential to outperform the best single-ISA heterogeneous architecture by as much as 21%, with 23% energy savings and a reduction of 32% in Energy Delay Product. | |||||
| Comment: This paper addresses the design of heterogeneous chip multiprocessors that combine cores with different instruction set architectures (ISAs), rather than restricting heterogeneity to microarchitecture alone as in prior single-ISA designs. The authors argue that ISA diversity, specifically leveraging Thumb, x86-64, and Alpha, offers additional axes of specialization such as code density, dynamic instruction count, and register pressure, which can be exploited for performance and energy gains. Key contributions include an exhaustive design-space exploration methodology to identify optimal heterogeneous-ISA core configurations under varying power and area budgets, and a compiler and runtime framework enabling process migration across diverse ISAs within a unified address space. New concepts introduced include a multi-ISA compilation methodology using fat binaries with target-specific code sections and common intermediate representations, a common page table structure spanning 32-bit and 64-bit ISAs, long-mode emulation extensions for Thumb (LD64/ST64 instructions), and a reverse data-flow analysis technique for fast stack state transformation at equivalence points. The runtime system combines dynamic binary translation (based on QEMU) with program state transformers to support seamless migration between ISAs. Evaluation uses SPEC CPU2006 benchmarks compiled via LLVM/Clang and simulated with gem5, modeling cores after ARM Cortex A-15, Alpha 21264, and Intel Core-i7. Results show the heterogeneous-ISA CMP achieves migration-safety in 45% of basic blocks with limited steady-state performance degradation. Compared to the best single-ISA heterogeneous architecture, the proposed design achieves up to 21% higher performance, 23% energy savings, and a 32% reduction in energy-delay product. The paper concludes that ISA diversity is a valuable additional dimension for heterogeneous multicore design, complementing microarchitectural heterogeneity. | |||||
BibTeX:
@article{Venkat2014,
author = {Venkat, Ashish and Tullsen, Dean M.},
title = {Harnessing ISA diversity: design of a heterogeneous-ISA chip multiprocessor},
journal = {SIGARCH Computer Architecture News},
publisher = {ACM},
year = {2014},
volume = {42},
number = {3},
pages = {121–132},
url = {/docs/Venkat2014.pdf},
doi = {https://doi.org/10.1145/2678373.2665692}
}
|
|||||
| Reghenzani, F., Bhuiyan, A., Fornaciari, W. and Guo, Z. | A Multi-Level DPM Approach for Real-Time DAG Tasks in Heterogeneous Processors | 2021 | Proceedings of the 42nd IEEE Real-Time Systems Symposium (RTSS), pp. 14–26 | inproceedings | DOI URL |
| Abstract: The modeling and analysis of real-time applications focus on the worst-case scenario because of their strict timing requirements. However, many real-time embedded systems include critical applications requiring not only timing constraints but also other system limitations, such as energy consumption. In this paper, we study the energy-aware real-time scheduling of Directed Acyclic Graph (DAG) tasks. We integrate the Dynamic Power Management (DPM) policy to reduce the Worst-Case Energy Consumption (WCEC), which is an essential requirement for energy-constrained systems. Besides, we extend our analysis with tasks' probabilistic information to improve the Average-Case Energy Consumption (ACEC), which is, instead, a common non-functional requirement of embedded systems. To verify the benefits of our approach in terms of reduced energy consumption, we finally conduct an extensive simulation, followed by an experimental study on an Odroid-H2 board. Compared to the state-of-the-art solution, our approach is able to reduce the power consumption up to 32.1%. | |||||
| Comment: The paper addresses energy-aware real-time scheduling of Directed Acyclic Graph (DAG) tasks on heterogeneous multiprocessor platforms, focusing on both Worst-Case Energy Consumption (WCEC), critical for energy-constrained systems, and Average-Case Energy Consumption (ACEC), relevant for general embedded system efficiency. The authors integrate multi-level Dynamic Power Management (DPM), generalizing processor idle states through the concept of C-States, to reduce energy consumption while respecting real-time constraints. A key contribution is extending break-even time analysis to multiple C-States per processor, enabling optimal idle-state selection under switching overheads. The work further incorporates probabilistic execution time information, using probabilistic execution time (pET) matrices and operations such as convolution, to estimate ACEC more accurately than worst-case-only models. The paper proposes algorithms for task decomposition, processor allocation, and idle interval computation across DAG nodes, building on prior intra-task merging techniques. Uncertainty in probabilistic estimation is addressed via the Dvoretzky-Kiefer-Wolfowitz inequality to bound measurement error. The approach is validated through simulation and real hardware experiments on an Odroid-H2 board running a PREEMPT_RT-patched Linux kernel, using DVFS to emulate heterogeneous core speeds. Compared to a state-of-the-art baseline (Fed Guo WCEC-policy), the proposed ACEC-policy achieves about 12.8% average energy reduction over the paper's own WCEC-policy and 28.6% over the baseline. Results also illustrate that theoretical WCEC bounds are notably pessimistic compared to observed average-case energy consumption in experimental runs. | |||||
BibTeX:
@inproceedings{Reghenzani2021,
author = {Reghenzani, Federico and Bhuiyan, Ashikahmed and Fornaciari, William and Guo, Zhishan},
title = {A Multi-Level DPM Approach for Real-Time DAG Tasks in Heterogeneous Processors},
booktitle = {Proceedings of the 42nd IEEE Real-Time Systems Symposium (RTSS)},
publisher = {IEEE},
year = {2021},
pages = {14–26},
url = {/docs/Reghenzani2021.pdf},
doi = {https://doi.org/10.1109/RTSS52674.2021.00014}
}
|
|||||
| Barbalace, A., Lyerly, R., Jelesnianski, C., Carno, A., Chuang, H.-R., Legout, V. and Ravindran, B. | Breaking the Boundaries in Heterogeneous-ISA Datacenters | 2017 | SIGARCH Computer Architecture News Vol. 45(1), pp. 645–659 |
article | DOI URL |
| Abstract: Energy efficiency is one of the most important design considerations in running modern datacenters. Datacenter operating systems rely on software techniques such as execution migration to achieve energy efficiency across pools of machines. Execution migration is possible in datacenters today because they consist mainly of homogeneous-ISA machines. However, recent market trends indicate that alternate ISAs such as ARM and PowerPC are pushing into the datacenter, meaning current execution migration techniques are no longer applicable. How can execution migration be applied in future heterogeneous-ISA datacenters?In this work we present a compiler, runtime, and an operating system extension for enabling execution migration between heterogeneous-ISA servers. We present a new multi-ISA binary architecture and heterogeneous-OS containers for facilitating efficient migration of natively-compiled applications. We build and evaluate a prototype of our design and demonstrate energy savings of up to 66% for a workload running on an ARM and an x86 server interconnected by a high-speed network. | |||||
| Comment: The paper addresses the problem of enabling execution migration for energy efficiency in future datacenters that contain servers with different instruction set architectures (ISAs). Current datacenter operating systems rely on migration only across homogeneous-ISA machines. The authors present a compiler, runtime, and operating system extension that together allow transparent migration of natively-compiled applications between heterogeneous-ISA servers. A new multi-ISA binary architecture is introduced, which aligns data and code symbols at identical virtual addresses across different ISAs. The system uses compiler-inserted migration points to safely transform thread state, including stack rewriting, between architectures. A heterogeneous distributed shared memory service maintains memory consistency during and after migration. The prototype extends Popcorn Linux and uses LLVM, evaluated on an ARM and x86 server pair. Experimental results with NAS Parallel Benchmarks show that dynamic scheduling across ISAs can achieve energy savings of up to 66% compared to static placement. The measured overheads from stack transformation and migration are small enough to not limit frequent thread migrations. Overall, the work demonstrates the feasibility of heterogeneous-ISA execution migration for improving datacenter energy efficiency. | |||||
BibTeX:
@article{Barbalace2017,
author = {Barbalace, Antonio and Lyerly, Robert and Jelesnianski, Christopher and Carno, Anthony and Chuang, Ho-Ren and Legout, Vincent and Ravindran, Binoy},
title = {Breaking the Boundaries in Heterogeneous-ISA Datacenters},
journal = {SIGARCH Computer Architecture News},
publisher = {ACM},
year = {2017},
volume = {45},
number = {1},
pages = {645–659},
url = {/docs/Barbalace2017.pdf},
doi = {https://doi.org/10.1145/3093337.3037738}
}
|
|||||
| Wang, X., Yeoh, S., Lyerly, R., Olivier, P., Kim, S.-H. and Ravindran, B. | A Framework for Software Diversification with ISA Heterogeneity | 2020 | Proceedings of the 23rd International Symposium on Research in Attacks, Intrusions and Defenses (RAID), pp. 427–442 | inproceedings | URL |
| Abstract: Software diversification is one of the most effective ways to defeat memory corruption based attacks. Traditional software diversification such as code randomization techniques diversifies program memory layout and makes it difficult for attackers to pinpoint the precise location of a target vulnerability. Some recent work in the architecture community uses diverse ISA configurations to defeat code injection or code reuse attacks, showing that dynamically switching the ISA on which a program executes is a promising direction for future security systems. However, most of these work either remain in a simulation stage or require extra efforts to write the program. In this paper, we propose HeterSec, a framework to secure applications utilizing a heterogeneous ISA setup composed of real-world machines. HeterSec runs on top of commodity x86_64 and ARM64 machines and gives the process the illusion that it runs on a multi-ISA chip multiprocessor (CMP) machine. With HeterSec, a process can dynamically select its underlying ISA environment. Therefore, a protected process would be capable of hiding the instruction set on which it executed or detecting abnormal program behavior by comparing execution results step-by-step from multiple ISA-diversified instances. To demonstrate the effectiveness of such a software framework, we implemented HeterSec on Linux and showcased its deployability by running it on a pair of x86_64 and ARM64 servers, connected over InfiniBand. We then conducted two case studies with HeterSec. In the first case, we implemented a multi-ISA moving target defense (MTD) system, which introduces uncertainty at the instruction set level. In the second case, we implemented a multi-ISA-based multi-version execution (MVX) system. The evaluation results show that HeterSec brings security benefits through ISA diversification with a reasonable performance overhead. |
|||||
| Comment: This paper addresses the challenge of using software diversification across real heterogeneous instruction set architectures (ISAs) to defend against memory corruption attacks. Existing approaches either rely on simulation or require significant developer effort. The main contribution is HeterSec, a framework that allows a process to execute across x86_64 and ARM64 machines, giving the illusion of a single multi-ISA chip multiprocessor. HeterSec enables two security applications: moving target defense (MTD) via function-level ISA switching, and multi-variant execution (MVX) with lock-step comparison of execution results across ISAs. It introduces a virtual descriptor table and remote procedure calls to transparently share system resources across machines, overcoming cross-OS kernel state differences. The framework leverages the Popcorn compiler to embed ISA migration metadata into binaries. Evaluations show that MTD on Nginx and Redis incurs modest overhead, while kernel-based MVX on I/O-intensive workloads exhibits lower overhead than a ptrace-based prototype. HeterSec demonstrates the feasibility of exploiting ISA heterogeneity for practical software diversification on commodity hardware. | |||||
BibTeX:
@inproceedings{Wang2020,
author = {Xiaoguang Wang and SengMing Yeoh and Robert Lyerly and Pierre Olivier and Sang-Hoon Kim and Binoy Ravindran},
title = {A Framework for Software Diversification with ISA Heterogeneity},
booktitle = {Proceedings of the 23rd International Symposium on Research in Attacks, Intrusions and Defenses (RAID)},
publisher = {USENIX Association},
year = {2020},
pages = {427–442},
url = {/docs/Wang2020.pdf}
}
|
|||||