Ryohei Kobayashi received his Ph.D. in Engineering from Tokyo Institute of Technology in 2016. From April 2016 to September 2024, he was an Assistant Professor at the Center for Computational Sciences, University of Tsukuba. Since October 2024, he has been an Associate Professor at the Supercomputing Research Center, Institute of Integrated Research, Institute of Science Tokyo. He has also held a concurrent appointment at the RIKEN Center for Computational Science since July 2021 and has been affiliated with the High Performance Big Data Team since January 2026. He leads the Advanced Computing ACceleration (AC2) Laboratory. His research focuses on high-performance computing, accelerator systems including GPUs, FPGAs, DPUs/SmartNICs, and wafer-scale systems, data movement and communication optimization, and runtime systems for sparse and irregular computing. His honors include the Vision Co-Creation Initiative Award (Early-Career Faculty Category, 2026), the HPC in Asia Poster Award (ISC 2018), and the IEICE CPSY Young Presentation Award (2015). He has served in program roles such as Proceedings Chair for HPC Asia 2026 and Publicity Co-Chair for IEEE Cluster 2025. He is a member of ACM, IEEE/IEEE CS, IPSJ, and IEICE.
Taiga Kobayashi works on communication optimization for large-scale systems, with interests in FPGAs, DPUs, machine learning, and data compression. His current work studies communication-data compression on NVIDIA BlueField DPUs for multi-GPU LLM training, with an eye toward communication substrates for future sparse and irregular accelerator workloads.
Research Interests: FPGA / DPU / Machine Learning / Data Compression
Akimasa Watanuki works on accelerator-oriented performance optimization, with interests in GPUs, CUDA, memory hierarchy, data layout optimization, and heterogeneous computing. His current work studies GNN training on GPUs and wafer-scale systems, focusing on irregular memory access, memory hierarchy behavior, and training efficiency in graph workloads.
Research Interests: Accelerators / GPUs / CUDA / Memory hierarchy / Data layout optimization / Accelerating parallel applications / Accelerating machine learning OSS / Heterogeneous computing (CPU-GPU, CPU-GPU-FPGA)