AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

async-programming communication distributed-computing fused-kernel gemm gpgpu hip kernel-fusion ml multigpu rdma remote-memory-access rma rocm shmem symmetric-memory triton workgroup-specialization
31 Open Issues Need Help Last updated: Aug 4, 2026

Open Issues Need Help

View All on GitHub
good first issue help wanted examples iris

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
documentation help wanted iris gluon

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
bug help wanted iris

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
bug enhancement help wanted

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
enhancement help wanted examples iris

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
enhancement good first issue help wanted test core iris

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
enhancement help wanted benchmarks examples test core iris

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
good first issue help wanted test

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
good first issue help wanted test

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
good first issue help wanted test

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
good first issue help wanted test

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
good first issue help wanted test

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
good first issue help wanted test

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
enhancement good first issue help wanted core

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
enhancement good first issue help wanted

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
good first issue help wanted examples

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Implement an atomic XOR operation within the Iris Triton-based RMA library, mirroring the functionality of Triton's existing atomic operations. This involves adding a new function to the Iris API that performs an atomic logical XOR on a specified memory location.

Complexity: 4/5
good first issue help wanted core

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Implement an atomic_or function within the Iris Triton-based RMA library, mirroring the functionality of Triton's existing atomic operations. This involves adding the functionality to the Iris API, ensuring compatibility with existing code and testing its correctness and performance.

Complexity: 4/5
good first issue help wanted core

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Implement an atomic minimum operation (`atomic_min`) within the Iris Triton-based RMA library, mirroring the functionality provided by Triton's existing atomic operations. This involves extending Iris's API to include this new function and ensuring its correct and efficient implementation within the multi-GPU context.

Complexity: 4/5
good first issue help wanted core

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Implement an atomic_max function within the Iris Triton-based RMA library, mirroring the functionality of Triton's existing atomic operations. This involves adding a new function to the Iris API that performs an atomic maximum operation on a specified memory location.

Complexity: 4/5
good first issue help wanted core

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Implement an atomic_and function within the Iris Triton-based RMA library, mirroring the functionality of Triton's existing atomic operations. This involves adding a new function to the Iris API that performs an atomic logical AND operation on a specified memory location.

Complexity: 4/5
good first issue help wanted core

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization
bug good first issue

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Make the currently private Iris HIP module public, allowing users to directly access its convenience functions (like `get_wall_clock_rate`) instead of going through the main Iris API. This involves changing the module's visibility and potentially updating documentation.

Complexity: 2/5
enhancement good first issue help wanted

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Refactor the GEMM + All-Scatter example in the Iris project to remove an unnecessary inner loop and restructure the communication to handle the full GEMM tile in a single operation. This involves removing redundant code and potentially rewriting the communication logic for improved efficiency.

Complexity: 4/5
good first issue help wanted examples

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization

AI Summary: Implement a producer-consumer example in Triton using Iris for multi-GPU communication. The example should utilize separate HIP streams for producer and consumer kernels, atomic signaling for synchronization, and demonstrate per-block synchronization using flags in shared memory. The code should handle data transfer between two GPUs using `iris.store` and `iris.load`.

Complexity: 4/5
good first issue help wanted examples

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python
#async-programming#communication#distributed-computing#fused-kernel#gemm#gpgpu#hip#kernel-fusion#ml#multigpu#rdma#remote-memory-access#rma#rocm#shmem#symmetric-memory#triton#workgroup-specialization