The rapid expansion of artificial intelligence (AI) applications across cloud and edge environments is driving the need for highly efficient hardware capable of supporting increasingly demanding workloads. Among these, matrix multiplication – particularly General Matrix Multiplication