# What would be a good explanation of PyFR having a good utilization of GPU acceleration?

**URL:** <https://pyfr.discourse.group/t/what-would-be-a-good-explanation-of-pyfr-having-a-good-utilization-of-gpu-acceleration/247>\
**Category:** General\
**Created:** [25 July 2019 19:47 UTC](https://pyfr.discourse.group/t/what-would-be-a-good-explanation-of-pyfr-having-a-good-utilization-of-gpu-acceleration/247 "2019-07-25T19:47:05Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Junting\_Chen](https://avatars.discourse-cdn.com/v4/letter/j/ac8455/32.png) [@Junting\_Chen](https://pyfr.discourse.group/u/Junting_Chen)\
**Post date:** [25 July 2019 19:47 UTC](https://pyfr.discourse.group/t/what-would-be-a-good-explanation-of-pyfr-having-a-good-utilization-of-gpu-acceleration/247/1 "2019-07-25T19:47:05Z")

</div>

Hello all,

I am trying to find a proper reason for the statement of “PyFR has a good utilization of GPU acceleration technology” , or possibly refer to a paper. I understand that memory transfer between CPU and GPU is one of the big slowdowns of GPU computing.

How would you say PyFR as a modern code uses GPU resource better than codes with histories then modified to utilize GPU acceleration?

In the paper “pyfr an opensource frame work for solving advection-diffusion…” in 2014, it was said GEMM was optimized for large square matrices, where the constant operator in PyFR are small and square, and state matrices are short and fat. Is this improved?

Junting Chen

---

<div class="post-metadata">

**Author:** ![p.vincent](https://yyz2.discourse-cdn.com/free1/user_avatar/pyfr.discourse.group/p.vincent/32/225_2.png) [@p.vincent](https://pyfr.discourse.group/u/p.vincent)\
**Post date:** [25 July 2019 20:31 UTC](https://pyfr.discourse.group/t/what-would-be-a-good-explanation-of-pyfr-having-a-good-utilization-of-gpu-acceleration/247/2 "2019-07-25T20:31:29Z")

</div>

Hi Junting

> I am trying to find a proper reason for the statement of “PyFR has a good utilization of GPU acceleration technology” , or possibly refer to a paper. I understand that memory transfer between CPU and GPU is one of the big slowdowns of GPU computing.

Here is an example of a paper that looks at performance aspects:

[https://ieeexplore.ieee.org/document/7876999](https://ieeexplore.ieee.org/document/7876999)

> In the paper “pyfr an opensource frame work for solving advection-diffusion…” in 2014, it was said GEMM was optimized for large square matrices, where the constant operator in PyFR are small and square, and state matrices are short and fat. Is this improved?

GEMM can actually perform well in a range of scenarios. However, we have also developed technology for smaller/sparse matrices:

[https://www.sciencedirect.com/science/article/pii/S0010465515004506](https://www.sciencedirect.com/science/article/pii/S0010465515004506)

Peter
