# Running PyFR on multiple GPUs on HPC cluster

**URL:** <https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136>\
**Category:** Cases\
**Tags:** hpc\
**Created:** [11 July 2024 14:54 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136 "2024-07-11T14:54:29Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![pyfr\_starting\_man](https://avatars.discourse-cdn.com/v4/letter/p/d26b3c/32.png) [@pyfr\_starting\_man](https://pyfr.discourse.group/u/pyfr_starting_man)\
**Post date:** [11 July 2024 14:54 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/1 "2024-07-11T14:54:29Z")

</div>

I have read through similar discussions on this, and I have seen to change device-id to local-rank for [backend-cuda] and to change GPUs to “compute exclusive mode”. However, I am running PyFR on supercomputer HPC cluster and do not have any sort of sudo access to change to compute exclusive mode. If someone could help me with a general idea as to how to run PyFR on multiple GPUs on supercomputer cluster it would be greatly appreciated. Maybe the only way is to submit a batch job rather than an interactive desktop. Perhaps modifying the beginning of my job script for 1 GPU would do the trick?:

```bash
#SBATCH --job-name="PyFR_Simulation"
#SBATCH --time=04:00:00
#SBATCH --output=myscript2.out
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=48
#SBATCH --gpus-per-node=1
#SBATCH --account=#####
#SBATCH --mail-type=BEGIN,END,FAIL

```

---

<div class="post-metadata">

**Author:** ![WillT](https://avatars.discourse-cdn.com/v4/letter/w/bc79bd/32.png) [@WillT](https://pyfr.discourse.group/u/WillT)\
**Post date:** [11 July 2024 14:57 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/2 "2024-07-11T14:57:53Z")

</div>

I you want to run a job on a node exclusively with slurm, you normally add something like this:

```auto
#SBATCH --exclusive
#SBATCH --nodes={some integer}

```

Also you probably want to set

```auto
#SBATCH --ntasks-per-gpu=1

```

---

<div class="post-metadata">

**Author:** ![pyfr\_starting\_man](https://avatars.discourse-cdn.com/v4/letter/p/d26b3c/32.png) [@pyfr\_starting\_man](https://pyfr.discourse.group/u/pyfr_starting_man)\
**Post date:** [11 July 2024 15:01 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/3 "2024-07-11T15:01:18Z")

</div>

Thanks for the quick reply! I am in need of more GPUs, as I am running a large grid with only 1 GPU and getting a CUDA out of memory error. When I run it as a batch job and request multiple GPUs, do I set device-id = local-rank for [backend-cuda]?

Also, when partitioning, should I be partitioning the grid in the number of GPUs that I have or the number of cores? i.e. if I have 4 GPUs and 48 cores, would I partition my mesh in 48 or 4? I had always assumed to do number of cores and that is why I have ntasks-per-node set to 48.

---

<div class="post-metadata">

**Author:** ![WillT](https://avatars.discourse-cdn.com/v4/letter/w/bc79bd/32.png) [@WillT](https://pyfr.discourse.group/u/WillT)\
**Post date:** [11 July 2024 15:17 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/4 "2024-07-11T15:17:23Z")

</div>

The device ID will be system-dependent, but most of the time `local-rank` is correct.

When running on GPU, yes, you want one partition per GPU you intend to run on.

---

<div class="post-metadata">

**Author:** ![pyfr\_starting\_man](https://avatars.discourse-cdn.com/v4/letter/p/d26b3c/32.png) [@pyfr\_starting\_man](https://pyfr.discourse.group/u/pyfr_starting_man)\
**Post date:** [11 July 2024 16:13 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/5 "2024-07-11T16:13:53Z")

</div>

Thank you! Also, one quick unrelated question. When I was running some of PyFR’s test cases, I would sometimes get a nice progress bar that appeared below GPUFreq = control\_disabled that updated the sim’s progress and ETA. However, most of the time I do not see it. Is there a way to have this always show up?

---

<div class="post-metadata">

**Author:** ![WillT](https://avatars.discourse-cdn.com/v4/letter/w/bc79bd/32.png) [@WillT](https://pyfr.discourse.group/u/WillT)\
**Post date:** [12 July 2024 10:55 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/6 "2024-07-12T10:55:56Z")

</div>

The progress bar is controlled by the `-p` flag, ie:

```bash
$ pyfr -p run ...

```

If you have that flag enabled but aren’t seeing the progress bar, then it has something to do with your environment rather than PyFR.

---

<div class="post-metadata">

**Author:** ![pyfr\_starting\_man](https://avatars.discourse-cdn.com/v4/letter/p/d26b3c/32.png) [@pyfr\_starting\_man](https://pyfr.discourse.group/u/pyfr_starting_man)\
**Post date:** [12 July 2024 16:39 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/7 "2024-07-12T16:39:46Z")

</div>

Thanks, that worked. Also, when running with multiple GPUs like we discussed earlier, I am seeing that only 1 of my GPUs is being used with local-rank. Here is what NVIDIA-smi shows me during one of the runs:

 ![image](https://global.discourse-cdn.com/free1/uploads/pyfr/original/1X/cc7af8a386473f3a309099470ade2fefdce2107c.png)  
How can I get it to run with all 4 GPUs?

---

<div class="post-metadata">

**Author:** ![fdw](https://avatars.discourse-cdn.com/v4/letter/f/a587f6/32.png) [@fdw](https://pyfr.discourse.group/u/fdw)\
**Post date:** [12 July 2024 22:46 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/8 "2024-07-12T22:46:48Z")

</div>

Try setting `device-id = local-rank`.

Regards, Freddie.

---

<div class="post-metadata">

**Author:** ![pyfr\_starting\_man](https://avatars.discourse-cdn.com/v4/letter/p/d26b3c/32.png) [@pyfr\_starting\_man](https://pyfr.discourse.group/u/pyfr_starting_man)\
**Post date:** [12 July 2024 23:15 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/9 "2024-07-12T23:15:08Z")

</div>

Yes, that is what I have it set at as I mentioned. Any other advice as how I can get it to run with all 4 GPUs? I am on a supercomputer cluster, so I do not have any sort of sudo access or anything. Thank you!

---

<div class="post-metadata">

**Author:** ![fdw](https://avatars.discourse-cdn.com/v4/letter/f/a587f6/32.png) [@fdw](https://pyfr.discourse.group/u/fdw)\
**Post date:** [13 July 2024 11:42 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/10 "2024-07-13T11:42:58Z")

</div>

Okay, how does `device-id = 0` work?

Regards, Freddie.

---

<div class="post-metadata">

**Author:** ![p.vincent](https://yyz2.discourse-cdn.com/free1/user_avatar/pyfr.discourse.group/p.vincent/32/225_2.png) [@p.vincent](https://pyfr.discourse.group/u/p.vincent)\
**Post date:** [18 July 2024 02:32 UTC](https://pyfr.discourse.group/t/running-pyfr-on-multiple-gpus-on-hpc-cluster/1136/11 "2024-07-18T02:32:18Z")

</div>

What system are you trying to run on? I would try setting --gpus-per-task=1. This should give each rank a GPU, which it sees as device 0. And then set device-id=0 in the .ini file.
