Using objects for data storage in high-performance I/O
Full text
USING OBJECTS FOR STORAGE IN HIGH-PERFORMANCE I/O Adrian Jackson ([email protected]) Andy Turner EPCC *Nicolau Manubens ECMWF
Storage
•Lots of ways to store data on storage devices •Filesystems have two components: •Data storage •Indexing •Data stored in blocks •Chunks of data physically stored on hardware somewhere •Indexing is used to associate names with blocks •File names are the index •Files may consist of many blocks •Variable sized nature of files makes this a hard problem to solve Filesystems … … inodes blocks or sectors
•Metadata required to match data to storage •Files and directories •Expandable files requires management of blocks of storage •Can lead to wasted storage •Can lead to resource limits •Can require metadata operations that reduce “performance” •Files wrap data in ways that may be not related to how data is created or accessed •Locking/sharing can be hard to achieve efficiently •Can require kernel space operations •Performance is fine for slow storage •Requires bulk operations for high performance Challenges of filesytems
•General vs specific hardware constraints •GPU is a specialised hardware version of a CPU – provides higher performance for reduced semantic flexibility in programs •Newer hardware provides more flexibility •Something like direct addressability with reduced hardware overheads •Access patterns for creating data may be different to consuming data •Many processes working on the same “data structures” •Metadata and searching becoming increasingly important for scientific datasets •More emphasis on read often datasets •Optimising everything in a single system setup is hard •Contention on shared systems can be a big issue Challenges of I/O for applications
•BSP (bulk synchronous parallel) I/O •Required for high performance on storage •Not necessarily what you would “naturally” do for an application •Workflows increasingly important Ideal I/O
Actual I/O ARCHER2 Lustre IO pattern analysis using Darshan
Actual I/O ARCHER2 Lustre IO pattern analysis using Darshan
•Another way to split up the storage hardware that is available… •…and manage the data being stored •Databases is one possible option •Requires a different model for shared access systems than is traditionally used for databases •Often not well designed for slice accessing or shared access to same dataset •ACID-like properties can be as restricting and POSIX approaches •Object storage an alternative approach •Deconstructed database Object storage as an alternative
Lustre vs DAOS
Ceph vs DAOS vs Lustre
Storage comparison – 16 storage nodes (or 16 + 1) GCP
Ceph vs Lustre vs DAOS: Small objects
•Testing under “production” conditions is important •Redundancy important for production operations •Even if backup not common •Potential to move away from redundancy if sites/groups/users want •Benefits of configurable redundancy •Two types of redundancy •Erasure coding •Replication DAOS under operational conditions
Nothing vs Erasure Coding (2+1) - IOR Write Read Nothing EC 2+1
Nothing vs Redundancy (factor 2) - IOR Write Read Nothing 2 x Replication
Erasure vs Redundancy Write Read 2 x Replication EC 2+1
2+1 Erasure coding
Access interfaces POSIX I/O / “Files” FUSE & Interception S3 Radosg w Block / NVMeoF SPDK DAOS bdev Pytho n pydaos Hadoop Connecto r MPI-IO DAOS ROMIO HDF5 DAOS VOL SEG Y FD B ROO T DA Q libdfs (Parallel Filesystem) libdaos (key-value-array interface) AI/Analytics/Scientific Workflow GPGPU CPU Compute Instances 1. Userspace DFS library with API like POSIX ○Require application changes ○Low latency & high concurrency ○No caching 2. DFUSE daemon to support POSIX API ○No application changes ○VFS mount point & high latency ○Caching by Linux kernel 3. DFUSE + Interception library ○No application changes ○2 flavors using LD_PRELOAD ○libioil ■(f)read/write interception ■Metadata via dfuse ○libpil4dfs ■Data & metadata interception ■Aim at delivering same performance as #1 w/o any application change ■Mmap & binary execution via fuse DFS - DAOS Filesystem (libdfs) DAOS Library (libdaos) Interception Library libpil4dfs libioil Application/Framework dfuse Single process address space Kernel bypass DAOS Storage Engine RPC RDMA System calls Linux Kernel Data & metadata Data 1 3b 3a 3 2 1 3a 3b 2
•For a specific benchmark run configured with contention across processes on indexing Key-Values: •20 GiB/s write •13 GiB/s read •Tweaking the benchmark configuration to have all processes operate on a separate Key-Values: •35 GiB/s write •68 GiB/s read •This may not be trivial or possible for all applications, but if design can support it then this improves performance Approach/recommendations: Key-Value contention
•Avoid communications on/with the server where possible •Cache objects locally in DRAM if possible •Use daos_array_open_with_attr to avoid daos_array_create calls •Only supported for DAOS_OT_ARRAY_BYTE, not for DAOS_OT_ARRAY •Warning: the cell size and chunk size attributes need to be provided consistently on any future daos_array_open_with_attr to avoid data corruption •daos_array_get_size calls can be expensive •Can store array size in our indexing Key-Values •Can manually calculate •Also possible to infer the size by reading with overallocation: •use DAOS_OT_ARRAY_BYTE, over-allocate the read buffer, and read without querying the size. The actual read size (short_read) will be returned •daos_cont_alloc_oids is expensive, call it just once per writer process •Required to generate object ideas to use in calls but can generate many at one Approach/recommendations
•Creating several containers (starting at ~300) in a DAOS pool reduces performance •Opening the same container from all processes is expensive •this happens even if only a few containers exist in the DAOS pool •e.g. out of 20 seconds taken by a process to write 2000 fields, 1.5 seconds were spent just to open one container •we observed this starting at ~200 parallel processes •Sharing handles using MPI is the way to fix this •Opening more than one container per process is very expensive •e.g. out of 30 seconds taken by a process to read 2000 fields, 6 seconds were spent just to open two containers Approach/recommendation
•daos_key_value_list is expensive •daos_array_open_with_attrs, daos_kv_open and daos_array_generate_oid are very cheap (no RPC) •Normal daos_array_open is expensive •daos_cont_alloc_oids is expensive •daos_kv_put and _get are generally cheap •Value size impacts this •daos_obj_close, daos_cont_close and daos_pool_disconnect are cheap •Server configuration to use available networks/sockets/etc… important for performance •Just like any storage system or application Approach/recommendations
Ceph
Designing for object stores
type, public, bind(c) :: daos_array_stbuf_t integer (kind=daos_size_t) :: st_size integer (kind=daos_epoch_t) :: st_max_epoch end type daos_array_stbuf_t interface integer(kind=c_int) function daos_array_create(coh, oid, th, cell_size, chunk_size, oh, ev) bind(c,name="daos_array_create") import :: c_int import :: daos_handle_t import :: daos_obj_id_t import :: daos_size_t import :: daos_event_t type(daos_handle_t), value, intent(in) :: coh type(daos_obj_id_t), value, intent(in) :: oid type(daos_handle_t), value, intent(in) :: th integer(kind=daos_size_t), value, intent(in) ::cell_size integer(kind=daos_size_t), value, intent(in) ::chunk_size type(daos_handle_t), intent(inout) :: oh type(daos_event_t), intent(inout) :: ev end function daos_array_create Nascent Fortran Interface
•DAOS has an MPI-I/O implementation •ROMIO in mpich 3.4.2 •Creates DAOS objects from MPI-I/O function calls… •…but does not target Array objects •Working on creating array objects from MPI-I/O specifications •Automatically mapping object sizes and extents, I/O sizes and extents, from MPI datatype and MPI-I/O pattern for applications •Very much a WiP at the moment •Simplifies porting “traditional” parallel I/O to DAOS more efficiently •Allows starting to tweak the setup but still staying with MPI-I/O MPI I/O DAOS automated object creation
•Object storage can provide high performance •DAOS: 90+ GB/s per server is possible •Hardware and configuration dependent, just like all I/O •Built in replication and redundancy under your/user control •Different interfaces available •Filesystem for zero cost porting •Simple file like access for slightly improved performance at little effort •Programming APIs for full functionality •Object store interface enables changing I/O granularity/patterns for bigger benefits •High performance when moving to more “realistic” deployment approaches •NVMe, replications/erasure coding •Provided interfaces are generally performant •But porting to the libdaos layer has benefits/potential, with an associated development cost Summary