# How to run Multi-GPU per node node with PyFR

**URL:** <https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139>\
**Category:** General\
**Created:** [28 February 2017 06:15 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139 "2017-02-28T06:15:20Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![AlmostSurelyRob](https://yyz2.discourse-cdn.com/free1/user_avatar/pyfr.discourse.group/almostsurelyrob/32/446_2.png) [@AlmostSurelyRob](https://pyfr.discourse.group/u/AlmostSurelyRob)\
**Post date:** [28 February 2017 06:15 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/1 "2017-02-28T06:15:20Z")

</div>

Dear All,

I just installed PyFR on our GPU cluster - what a breeze. I can run your examples and tutorials with CUDA backend and I can see them being offloaded to the GPU. Metis also appears to cooperate.

My question is how to configure multiple GPUs on a single node. Is there a way to make sure that say four processes will run on four different GPUs exclusively?

Also, would be possible to share a larger case as an example to actually flood the GPUs with work. I am more than willing to give PyFR a try but it will take a while for me to develop case setting skills.

Many thanks,  
Robert

---

<div class="post-metadata">

**Author:** ![fdw](https://avatars.discourse-cdn.com/v4/letter/f/a587f6/32.png) [@fdw](https://pyfr.discourse.group/u/fdw)\
**Post date:** [28 February 2017 13:19 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/2 "2017-02-28T13:19:13Z")

</div>

Hi Robert,

The easiest way is to put the GPUs into compute exclusive mode. This can be accomplished using the nvidia-smi tool and ensures that only one process can use a GPU at any given time.

Alternatively, you can set

[backend-cuda]  
device-id = local-rank

which will assign the first MPI rank on the node to the first CUDA device, and so on and so forth. This approach does not require that the GPUs be in compute exclusive mode.

Regards, Freddie.

---

<div class="post-metadata">

**Author:** ![arvind\_iyer](https://avatars.discourse-cdn.com/v4/letter/a/a6a055/32.png) [@arvind\_iyer](https://pyfr.discourse.group/u/arvind_iyer)\
**Post date:** [28 February 2017 14:25 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/3 "2017-02-28T14:25:45Z")

</div>

Hi,

Running separate ranks on different GPUs can be done in two ways:

1. adding a section:

```auto
[backend-cuda]
device-id = local rank

```

see: [http://www.pyfr.org/user\_guide.php](http://www.pyfr.org/user_guide.php) for more details

1. running `nvidia-smi -c 1` as root.  
This will enforce one compute process per card.

2. To flood all GPUs/ increase the load, you can simply increase the polynomial order for the solution  
polynomial and/or add suitable anti-aliasing in any test case.

Regards  
Arvind

---

<div class="post-metadata">

**Author:** ![AlmostSurelyRob](https://yyz2.discourse-cdn.com/free1/user_avatar/pyfr.discourse.group/almostsurelyrob/32/446_2.png) [@AlmostSurelyRob](https://pyfr.discourse.group/u/AlmostSurelyRob)\
**Post date:** [28 February 2017 22:04 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/4 "2017-02-28T22:04:02Z")

</div>

Thanks Freddie and Arvin,

Somebody just made it too easy… Thanks. With the backend option I can run across nodes and with multiple GPUs per node.

My last question was a little cheeky I admit. I was curious if the cases from Vermeire, Witherden and Vincent (JCP 2017) paper are available. I can see the zip file among supplementary material but was unable to download it from Elsevier page even though the journal is open access. I’d like to use these cases mainly to test and benchmark two GPU clusters.

Robert

---

<div class="post-metadata">

**Author:** ![p.vincent](https://yyz2.discourse-cdn.com/free1/user_avatar/pyfr.discourse.group/p.vincent/32/225_2.png) [@p.vincent](https://pyfr.discourse.group/u/p.vincent)\
**Post date:** [28 February 2017 22:33 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/5 "2017-02-28T22:33:51Z")

</div>

Hi Robert,

Thanks for your interest in PyFR!

> My last question was a little cheeky I admit. I was curious if the cases from Vermeire, Witherden and Vincent (JCP 2017) paper are available. I can see the zip file among supplementary material but was unable to download it from Elsevier page even though the journal is open access. I’d like to use these cases mainly to test and benchmark two GPU clusters.

I was just going to suggest this. But as you say, it seems like their site is down. I’ll email Elsevier now. If its not back up tomorrow I’ll send over the files directly.

Cheers

Peter

---

<div class="post-metadata">

**Author:** ![bvermeir](https://avatars.discourse-cdn.com/v4/letter/b/a698b9/32.png) [@bvermeir](https://pyfr.discourse.group/u/bvermeir)\
**Post date:** [1 March 2017 00:12 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/6 "2017-03-01T00:12:39Z")

</div>

Hi Robert,

I just heard back from the publisher and they were doing site maintenance earlier, which is why the files were unavailable. You should be able to download the supplementary material now, but let us know if still have any problems accessing them.

Cheers,

---

<div class="post-metadata">

**Author:** ![AlmostSurelyRob](https://yyz2.discourse-cdn.com/free1/user_avatar/pyfr.discourse.group/almostsurelyrob/32/446_2.png) [@AlmostSurelyRob](https://pyfr.discourse.group/u/AlmostSurelyRob)\
**Post date:** [1 March 2017 00:16 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/7 "2017-03-01T00:16:54Z")

</div>

Thanks for letting me know. I downloaded them already and got them to produce  
output on our cluster. I am focused on the isentropic vortex for now. There  
were a few minor changes I had to do in ini such as fixing "tend" and  
restructuring the output section.

Quick question. I am considering creating a github repository with the up-to  
date versions for these examples for 1.5.0. Is that ok with you?

Best wishes,  
Robert

---

<div class="post-metadata">

**Author:** ![AlmostSurelyRob](https://yyz2.discourse-cdn.com/free1/user_avatar/pyfr.discourse.group/almostsurelyrob/32/446_2.png) [@AlmostSurelyRob](https://pyfr.discourse.group/u/AlmostSurelyRob)\
**Post date:** [13 March 2017 09:51 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/8 "2017-03-13T09:51:55Z")

</div>

Peter and Brian,

I am just forwarding this message as I thought it should have gone to  
the list.

Thanks. I think you answered for now all my initial questions about PyFR. Also,  
I finally sat down and compiled the SD7003 results from two GPU clusters to  
which I have access and I thought I’ll share. Let me summarise a few things.

- It seems to me that binding to sockets matters on IBM processors. Initially,  
I was binding everything to socket 1 and I was getting either freezing  
behaviour or no scaling.
- With core and 2 per socket I seem to be getting stable behaviour and  
scaling up to 16 nodes.
- I wasn’t able to run with `cuda-aware` switch despite compiling OpenMPI with  
CUDA. This is something I am still working on.

There still may be issues with clusters or software stack as both are in pretty  
early days in terms of operation, but preliminary results look good. Let me  
know what you think.

This is just a dry run of your code just to prove that it can work in  
principle. I’m interested in looking into some details and pushing the  
development.

This is the repository I created:

> **[robertsawko/PyFR-bench](https://github.com/robertsawko/PyFR-bench)**
>
> Benchmarks for PyFR. Contribute to robertsawko/PyFR-bench development by creating an account on GitHub.

Best wishes,  
Robert

[sd7003-absolute.pdf](https://pyfr.discourse.group/uploads/short-url/szIubwwlZ4bPBOxlzfCIfC9ARFt.pdf) (50.3 KB)

[sd7003-efficiency.pdf](https://pyfr.discourse.group/uploads/short-url/pqehFzsy2lvhpB4We970qVtVjY7.pdf) (84.2 KB)

[sd7003-speedup.pdf](https://pyfr.discourse.group/uploads/short-url/cyTDVu5VF5mfOsLrGx4Z57VZzjx.pdf) (50.2 KB)

---

<div class="post-metadata">

**Author:** ![nnunn](https://avatars.discourse-cdn.com/v4/letter/n/c89c15/32.png) [@nnunn](https://pyfr.discourse.group/u/nnunn)\
**Post date:** [28 March 2017 14:47 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/9 "2017-03-28T14:47:42Z")

</div>

Hi Robert - thanks for setting up the SD7003 example. I just tried a short MPI run (metis + Windows 7) on my 4 Kepler Titans. Runs fine, keeping all four GPUs steady at over 95% utilization.

To PyFR team: can you post the gmsh geometry file used for extruding the SD7003 profile? I’d like to try running a 2D version over that profile.

many thanks for making all this available,  
Nigel

---

<div class="post-metadata">

**Author:** ![Robert\_Sawko1](https://avatars.discourse-cdn.com/v4/letter/r/9de053/32.png) [@Robert\_Sawko1](https://pyfr.discourse.group/u/Robert_Sawko1)\
**Post date:** [28 March 2017 15:22 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/10 "2017-03-28T15:22:20Z")

</div>

Hi Nigel,

No problem. I've really just redone things that PyFR developers put together  
with their paper. It's really great when people do reproducible work-flows.  
After this Wednesday I will be back to HPC work and will extend the benchmark  
with Taylor-Green case and any updates for the new version of PyFR.

Also, I am interested to do some visualisation for the SD7003 case. Is there  
any way I could get hold of the original geometry so that I can show for  
instance Q-criterion + wing shape. The default vtu output only shows the  
internal fluid field and I can't select components like BCs or surfaces.

Best wishes,  
Robert

---

<div class="post-metadata">

**Author:** ![bvermeir](https://avatars.discourse-cdn.com/v4/letter/b/a698b9/32.png) [@bvermeir](https://pyfr.discourse.group/u/bvermeir)\
**Post date:** [20 April 2017 13:29 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/11 "2017-04-20T13:29:24Z")

</div>

Hi Robert,

Please find attached a .vtu file of the SD7003 wall. You may need to rotate it in paraview depending on what angle of attack you are running.

Hi Nigel,

Please find attached a Gmsh .geo file for the SD7003 airfoil. Please note that this will not generate an identical mesh to the one used in the paper. Even an identical .geo file will generate a different mesh depending on your version of Gmsh, settings, etc. If you want to reproduce the paper results just use the .msh file provided in the supplementary material. However, if you want to run 2D simulations you should be able to get a case working from this .geo file.

Cheers,

[sd7003.geo](https://pyfr.discourse.group/uploads/short-url/AfwWfkJ1T60KevO3nbgyO5KqWkY.geo) (10.1 KB)

[wall.vtu](https://pyfr.discourse.group/uploads/short-url/a45MR5Zljx5IJMOGMvLer51DRZY.vtu) (431 KB)

---

<div class="post-metadata">

**Author:** ![nnunn](https://avatars.discourse-cdn.com/v4/letter/n/c89c15/32.png) [@nnunn](https://pyfr.discourse.group/u/nnunn)\
**Post date:** [20 April 2017 16:20 UTC](https://pyfr.discourse.group/t/how-to-run-multi-gpu-per-node-node-with-pyfr/139/12 "2017-04-20T16:20:02Z")

</div>

Hi Brian - thanks for sending the .geo file. What a great example for driving gmsh!  
Those 2d nodes defining the wall profile are exactly what I needed.  
with thanks - Nigel
