# Cuda backend parameters

**URL:** <https://pyfr.discourse.group/t/cuda-backend-parameters/262>\
**Category:** General\
**Created:** [6 January 2020 17:34 UTC](https://pyfr.discourse.group/t/cuda-backend-parameters/262 "2020-01-06T17:34:40Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Junting\_Chen](https://avatars.discourse-cdn.com/v4/letter/j/ac8455/32.png) [@Junting\_Chen](https://pyfr.discourse.group/u/Junting_Chen)\
**Post date:** [6 January 2020 17:34 UTC](https://pyfr.discourse.group/t/cuda-backend-parameters/262/1 "2020-01-06T17:34:40Z")

</div>

Hello,

I am wondering if someone can provide a bit more descriptions on these parameters to optimize performance.

As far as I know, when using multiple GPUs, I had to select local-rank for device-id and cuda-aware for mpi-type. When exactly should i be using round-robin and local-rank? And when should i be using standard or cuda-aware?

How would you select GiMMiK cutoff? How does it affect accuracy / performance?

I believe block-1d and block-2d are determined by GPU’s specification. I am not very familiar with Cuda. Please someone can elaborate a bit. For example I am running pyfr with two Tesla k80s in parallel, what’s the block size for 1d and 2d pointswise kernels?

Parameterises the CUDA backend with

1. `device-id` — method for selecting which device(s) to run on:

2. `gimmik-max-nnz` — cutoff for GiMMiK in terms of the number of non-zero entires in a constant matrix:

3. `mpi-type` — type of MPI library that is being used:

4. `block-1d` — block size for one dimensional pointwise kernels:

5. `block-2d` — block size for two dimensional pointwise kernels:

Thanks a lot!

Junting Chen

---

<div class="post-metadata">

**Author:** ![fdw](https://avatars.discourse-cdn.com/v4/letter/f/a587f6/32.png) [@fdw](https://pyfr.discourse.group/u/fdw)\
**Post date:** [6 January 2020 22:18 UTC](https://pyfr.discourse.group/t/cuda-backend-parameters/262/2 "2020-01-06T22:18:24Z")

</div>

Hi Junting,

> As far as I know, when using multiple GPUs, I had to select local-rank  
> for device-id and cuda-aware for mpi-type. When exactly should i be  
> using round-robin and local-rank? And when should i be using standard or  
> cuda-aware?

If the GPUs in your system are in compute exclusive mode then  
round-robin is probably what you want. Otherwise, opt for local-rank.  
So long as each rank gets its own GPU there should be no impact on  
performance.

In terms of the mpi-type this depends heavily on the hardware you're  
running on and the MPI library you're using. If your MPI library is  
CUDA aware then setting mpi-type = cuda-aware can improve performance.

> How would you select GiMMiK cutoff? How does it affect accuracy /  
> performance?

Some experimentation is needed here as the optimal value depends on the  
element types you're using, if anti-aliasing is enabled, and the CPU  
that you are running on.

> I believe block-1d and block-2d are determined by GPU's specification. I  
> am not very familiar with Cuda. Please someone can elaborate a bit. For  
> example I am running pyfr with two Tesla k80s in parallel, what's the  
> block size for 1d and 2d pointswise kernels?

You should seldom need to modify either of these two values. On some  
pathological meshes reducing block-1d can improve performance, but not  
by a lot.

Regards, Freddie.

---

<div class="post-metadata">

**Author:** ![Junting\_Chen](https://avatars.discourse-cdn.com/v4/letter/j/ac8455/32.png) [@Junting\_Chen](https://pyfr.discourse.group/u/Junting_Chen)\
**Post date:** [7 January 2020 17:40 UTC](https://pyfr.discourse.group/t/cuda-backend-parameters/262/3 "2020-01-07T17:40:47Z")

</div>

Thanks Freddie,

So when starting a run, do you usually play with the GiMMiK cutoff a bit to find the most optimized value (does it influence the performance significantly / worth the effort of finding the optimized value)? What’s the range of this value? Is a power of 2 (example uses 512) somewhat beneficial?

Junting

---

<div class="post-metadata">

**Author:** ![fdw](https://avatars.discourse-cdn.com/v4/letter/f/a587f6/32.png) [@fdw](https://pyfr.discourse.group/u/fdw)\
**Post date:** [8 January 2020 14:32 UTC](https://pyfr.discourse.group/t/cuda-backend-parameters/262/4 "2020-01-08T14:32:01Z")

</div>

Hi Junting,

The values I would try are 0 (disables GiMMiK), 512 (the default), and  
8192. There is nothing special about the number being a power of two.

Regards, Freddie.

---

<div class="post-metadata">

**Author:** ![WillT](https://avatars.discourse-cdn.com/v4/letter/w/bc79bd/32.png) [@WillT](https://pyfr.discourse.group/u/WillT)\
**Post date:** [23 May 2021 21:50 UTC](https://pyfr.discourse.group/t/cuda-backend-parameters/262/5 "2021-05-23T21:50:04Z")

</div>

6 posts were split to a new topic: [Cuda backend error ‘CudaOutOfMemory’ ‘CudaInvaildDevice’](https://pyfr.discourse.group/t/cuda-backend-error-cudaoutofmemory-cudainvailddevice/401)

---

<div class="post-metadata">

**Author:** ![WillT](https://avatars.discourse-cdn.com/v4/letter/w/bc79bd/32.png) [@WillT](https://pyfr.discourse.group/u/WillT)\
**Post date:** [23 May 2021 21:41 UTC](https://pyfr.discourse.group/t/cuda-backend-parameters/262/11 "2021-05-23T21:41:45Z")

</div>


