scieee Science in your language
[en] (orig)

HyperQueue: overcoming limitations of HPC job managers

Read accessible full text

HyperQueue: overcoming limitations of HPC job managers

Author: Böhm, Stanislav
Year: 2021
Source: https://dspace.vsb.cz/bitstreams/7ab8103d-512f-4ba8-8f30-8c287b2f9599/download
Hype Queue: O e coming Limi a ions o HPC Job Manage s
S anisla Böhm
s anisla [email p o ec ed]
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
Jakub Be ánek
jakub[email p o ec ed]
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
Voj ěch Cima
oj ech.cima@ sb.cz
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
Roman Macháček
oman.machacek@ sb.cz
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
Vyomkesh Jha
jha0007@ sb.cz
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
Al éd Kočí
[email p o ec ed]
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
B anisla Jansík
b anisla .jansik@ sb.cz
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
Jan Ma ino ič
jan.ma ino ic@ sb.cz
IT4Inno a ions, VSB – Technical
Uni e si y o Os a a
Czech Republic
ABSTRACT
In ecen yea s, HPC wo kloads and communi ies ha e unde gone
subs an ial pa adigm shi s. The e is an inc easing amoun o use s
ha wan o le e age HPC clus e s o execu e many simple and
emba assingly pa allel asks as easily as possible. Howe e , due
o he limi a ions o adi ional HPC job manage s, hese use s
mus o en eso o manual agg ega ion o asks in o a smalle
numbe o jobs o educe job manage o e head. This app oach
is bo h labou -in ensi e and ine icien , as i lacks dynamic load
balancing equi ed o ully u ilize compu a ional nodes wi h ens
o hund eds o co es. We in oduce Hype Queue - a ask scheduling
un ime ha can execu e a la ge amoun o asks on op o an
HPC job manage by au oma ically agg ega ing asks in o jobs and
dynamically load balancing hem ac oss all alloca ed nodes and
CPU co es. Hype Queue is an open-sou ce ool ha is designed o
ease o use and deploymen .
KEYWORDS
scheduling, ask managemen , esou ce managemen , clus e
1 INTRODUCTION
Mode n HPC sys ems con ain a la ge numbe o compu a ional
nodes, wi h each node ha ing ens o hund eds o CPU co es. The e
is an inc easing numbe o HPC use s ha wan o le e age all his
compu a ional powe wi h qui e simple wo kloads, o example by
unning a la ge numbe o single-node, ela i ely sho -li ed asks
in an emba assingly pa allel manne [1].
While HPC sys ems a e p epa ed o almos a bi a ily complex
dis ibu ed compu a ions, i can be su p isingly di icul o execu e
a la ge numbe o such simple asks on hem e icien ly. HPC job
manage s such as Slu m [
2
] o PBS a e almos ubiqui ously used
o manage compu a ional esou ces o HPC clus e s. The mos
s aigh o wa d app oach in his scena io is hus o simply map
each ask di ec ly o a single job.
Howe e , his is usually in easible, because he manage in o-
duces non i ial o e head and hus se e ely limi s he amoun o
jobs ha can be submi ed by a single use . The manage is also
usually con igu ed in a way ha does no allow alloca ing sub-node
esou ces ( unning mul iple jobs on a single node concu en ly),
ei he o secu i y o pe o mance s abili y easons. The e o e, i
he asks only use a ew co es, hey canno ully u ilize he p o ided
compu a ional nodes. These limi a ions can pose a ba ie o en y
o new HPC use s, since hey canno submi asks o he clus e in
a s aigh o wa d way.
To o e come hese limi a ions, we a e in oducing Hype Queue,
a ask execu ion un ime ha can e icien ly execu e many single-
node asks on op o a Slu m/PBS clus e . Use s can submi a la ge
amoun o asks in o Hype Queue in a simple way; i will hen
au oma ically agg ega e asks in o a smalle amoun o jobs. Once
he jobs a e s a ed, i will dis ibu e and dynamically load balance
he asks ac oss all alloca ed jobs, nodes and co es o e icien ly
u ilize a ailable esou ces.
Hype Queue is an open-sou ce Rus p ojec eleased unde he
MIT licence: h ps://gi hub.com/I 4inno a ions/hype queue.
2 RELATED WORK
The e a e se e al o he app oaches ha can be used ins ead o
c ea ing a sepa a e job o each ask. One al e na i e is o un an
addi ional schedule on op o ("ou side") he in e nal job manage .
Snakemake [
3
] can agg ega e mul iple asks in o g oups. This ap-
p oach educes he numbe o submi ed jobs and hus dec eases
he o e head in oduced by he job manage . Ne e heless, his
agg ega ion is pe o med eage ly, be o e he asks e en s a o
execu e. The e o e, he asks canno be load balanced dynamically
in esponse o ac ual node and co e u iliza ion.
Ano he op ion is o s a a scheduling un ime wi hin ("inside")
aSlu m/PBS job o enable load balancing o asks independen ly o
he ou e job manage . The main disad an age o his app oach is
Böhm e al.
ha he use has o manually agg ega e asks in o jobs o a oid c e-
a ing oo many o hem. Such agg ega ion is no always i ial and
i does no enable ask load balancing ac oss di e en jobs, which
may cause ine icien node u iliza ion. As an example, Dask [
5
] can
be used in his way.
To achie e bo h load balancing and au oma ic ask agg ega ion,
a scheduling sys em has o un bo h ou side he job manage as
well as inside he submi ed jobs. Me lin [
4
] uses his app oach, bu
in oduces se e al o he limi a ions. I s load balancing is limi ed,
because i equi es use s o speci y queues wi h p ede e mined
concu ency and also does no allow speci ying ask esou ce e-
qui emen s. I is also non i ial o deploy, since i equi es se ing
up se e al complex dependencies.
We a e no awa e o any exis ing ool ha au oma ically agg e-
ga es asks and also ully dynamically load balances hem while
being easy o deploy and use. We conside hese p ope ies c ucial
o enabling use - iendly execu ion o wo k lows con aining many
emba assingly pa allel asks on HPC clus e s.
3 HYPERQUEUE
Hype Queue is a ask scheduling un ime designed o e icien ly exe-
cu e a la ge numbe o emba assingly pa allel asks on a Slu m/PBS
clus e in a simple way. We lis he key ea u es o Hype Queue in
he ollowing lis .
Task Agg ega ion Hype Queue does no equi e any manual
ask agg ega ion. Use s simply submi asks in o Hype Queue and
i akes ca e o dis ibu ing hem o compu a ional esou ces (nodes
and co es) alloca ed by Slu m/PBS. Hype Queue can ei he use
compu a ional esou ces (jobs) alloca ed ex e nally, o i can submi
jobs au oma ically, as needed by he cu en wo kload.
Load balancing Hype Queue uses a wo k-s ealing schedule
ha suppo s ask esou ce equi emen s (how many co es does a
ask need), ask p io i ies, co e pinning and is NUMA-awa e. The
schedule is dynamic wi h espec o compu a ional esou ces; new
wo ke s may be added o emo ed du ing wo kload compu a ion
and new asks may be submi ed a any ime.
Scaling Hype Queue suppo s ask a ays and also ou pu s eam-
ing, which educes p essu e on ne wo k ilesys ems commonly
p esen on HPC clus e s. The a e age un ime o e head is a ound
100𝜇𝑠 pe ask, which allows p ocessing o ine-g ained asks.
Deploymen Hype Queue is dis ibu ed as a single bina y ha
depends only on
libc
, does no need any hi d-pa y se ices, and
does no equi e any ele a ed p i ileges; i is hus i ial o deploy.
3.1 A chi ec u e
The a chi ec u e o Hype Queue is depic ed in Figu e 1. The cen al
componen is he se e , which manages wo ke s and schedules
asks on o hem. I is designed o be a long-li ed p ocess which can
be deployed e.g. on a login node. Use s can connec o i in o de o
submi asks o obse e he s a us o compu a ion.
The wo ke componen uns on he compu a ional nodes o a
clus e , inside a Slu m o a PBS job. I s main pu pose is o execu e
asks assigned o i by he se e . The wo ke s communica e wi h
he se e o e he ne wo k; hey mus be able o ini ia e a TCP/IP
connec ion o he se e .
HQ
Se e
SLURM
/
PBS
HQ
Wo ke
$ hq submi echo "Hello wo ld!"
$ hq submi --cpus 16 my-compu a ion.sh
HQ
Wo ke echo "Hello wo ld!"
Task [1 cpu]
Node
HQ
Wo ke
HQ
Wo ke
Node
my-compu a ion.sh
Task [16 cpus]
Op ional: s eaming ou pu s
Figu e 1: Hype Queue a chi ec u e
3.2 Usage example
(1) S a se e on a login node:
$ hq se e s a &
(2) Ask o a Slu m/PBS job and s a Hype Queue wo ke
•Ask o nodes in PBS:
$ qsub <qsub-a gs> -- hq wo ke s a
•Ask o nodes in SLURM:
$ sba ch <sba ch-a gs> --w ap "hq wo ke s a "
(3) Submi a ask:
$ hq submi my-compu a ion.sh
(4) Check he s a us o he compu a ion:
ACKNOWLEDGEMENT
This wo k was suppo ed by he LIGATE p ojec . This p ojec has
ecei ed unding om he Eu opean High-Pe o mance Compu ing
Join Unde aking (JU) unde g an ag eemen No 956137. The JU
ecei es suppo om he Eu opean Union’s Ho izon 2020 esea ch
and inno a ion p og amme and I aly, Sweden, Aus ia, he Czech
Republic, Swi ze land.
This wo k was suppo ed by he Minis y o Educa ion, You h
and Spo s o he Czech Republic h ough he e-INFRA CZ (ID:90140).
REFERENCES
[1]
Dong H. Ahn, Ned Bass, Albe Chu, Jim Ga lick, Ma k G ondona, S ephen He -
bein, Joseph Koning, Tapasya Pa ki, Thomas R. W. Scogland, Becky Sp ingmeye ,
and Michela Tau e . 2018. Flux: O e coming Scheduling Challenges o Exas-
cale Wo k lows. In 2018 IEEE/ACM Wo k lows in Suppo o La ge-Scale Science
(WORKS). 10–19. h ps://doi.o g/10.1109/WORKS.2018.00007
[2]
Mo is A. Je e, Andy B. Yoo, and Ma k G ondona. 2002. SLURM: Simple Linux U il-
i y o Resou ce Managemen . In In Lec u e No es in Compu e Science: P oceedings
o Job Scheduling S a egies o Pa allel P ocessing (JSSPP) 2003. Sp inge -Ve lag,
44–60.
[3]
Johannes Kös e and S en Rahmann. 2012. Snakemake—a scal-
able bioin o ma ics wo k low engine. Bioin o ma ics 28, 19
(08 2012), 2520–2522. h ps://doi.o g/10.1093/bioin o ma ics/
b s480 a Xi :h ps://academic.oup.com/bioin o ma ics/a icle-
pd /28/19/2520/819790/b s480.pd
[4]
Luc Pe e son, Rushil Ani udh, Ke in A hey, Benjamin Bay, Pee -Timo B eme ,
Vic Cas illo, F ancesco Na ale, Da id Fox, Jim Ga ney, Da id Hysom, Sam Jacobs,
Bha ya Kailkhu a, Joe Koning, Bogdan Kus owski, S e en Lange , Pe e Robinson,
Jessica Semle , B ian Spea s, Jaya aman J. Thiaga ajan, and Jae-Seung Yeom. 2019.
Me lin: Enabling Machine Lea ning-Ready HPC Ensembles.
[5]
Ma hew Rocklin. 2015. Dask: Pa allel Compu a ion wi h Blocked algo i hms and
Task Scheduling. In P oceedings o he 14 h Py hon in Science Con e ence, Ka h yn
Hu and James Be gs a (Eds.). 130 – 136.