DOLFINx 0.12.0.0
DOLFINx C++
Loading...
Searching...
No Matches
Scatterer< Container > Class Template Reference

A Scatterer supports the scattering and gathering of distributed data that is associated with a common::IndexMap, using MPI. More...

#include <Scatterer.h>

Public Types

using container_type = Container
 Container type used to store local and remote indices.

Public Member Functions

 Scatterer (const IndexMap &map)
 Create a scatterer for data with a layout described by an IndexMap.
template<class U>
 Scatterer (const Scatterer< U > &s)
 Cast-copy constructor.
 Scatterer (const Scatterer &scatterer)=default
 Scatterer (Scatterer &&scatterer)=default
 ~Scatterer ()=default
Scatterer & operator= (const Scatterer &scatterer)=delete
Scatterer & operator= (Scatterer &&scatterer)=default
template<typename T>
void scatter_fwd_begin_dtype (const T *send_buffer, T *recv_buffer, MPI_Datatype type, MPI_Request &request) const
 Start a non-blocking neighbourhood collective exchange of owned data with the ranks that ghost it.
template<typename T>
void scatter_fwd_begin (const T *send_buffer, T *recv_buffer, int bs, MPI_Request &request) const
 Start a non-blocking neighbourhood collective exchange of owned data with the ranks that ghost it.
void scatter_fwd_end (MPI_Request &request) const
 Complete a non-blocking MPI neighbourhood collective send.
template<typename T>
void scatter_rev_begin_dtype (const T *send_buffer, T *recv_buffer, MPI_Datatype type, MPI_Request &request) const
 Start a non-blocking neighbourhood collective exchange of ghost data with the owning ranks.
template<typename T>
void scatter_rev_begin (const T *send_buffer, T *recv_buffer, int bs, MPI_Request &request) const
 Start a non-blocking neighbourhood collective exchange of ghost data with the owning ranks.
void scatter_rev_end (MPI_Request &request) const
 Complete a non-blocking MPI neighbourhood collective send.
const container_type & local_indices_block () const noexcept
 Array of indices for packing/unpacking owned data to/from a send/receive buffer.
const container_type & remote_indices_block () const noexcept
 Array of indices for packing/unpacking ghost data to/from a send/receive buffer.

Detailed Description

template<class Container = std::vector<std::int32_t>>
class dolfinx::common::Scatterer< Container >

A Scatterer supports the scattering and gathering of distributed data that is associated with a common::IndexMap, using MPI.

Scatter and gather operations use MPI neighbourhood collectives.

The implementation is designed for sparse communication patterns, as is typical of patterns based on an IndexMap.

A Scatterer is stateless, i.e. it provides the required information and static data for a given parallel communication pattern but does not provide any communication caches or track the status of MPI requests. Callers of a Scatterer's members are responsible for managing buffer and MPI request handles.

Creating, copying and destroying a Scatterer are collective, since they create, duplicate and free MPI communicators. Move construction is not collective, but move assignment is, since it frees the communicators held by the assignment target.

A forward scatter sends data associated with owned/local indices to the ranks that ghost them; a reverse scatter sends ghost data back to the owning ranks, to be accumulated into the owned data. Both use the same two-step begin/end pattern, splitting the non-blocking MPI call from completion so that unrelated work can be done while communication is in flight. A round trip for a forward scatter with block size 1, where x holds the owned data and x_ghost the ghost data:

Scatterer sc(map);
std::vector<std::int64_t> send_buffer(sc.local_indices_block().size());
{
auto& idx = sc.local_indices_block();
for (std::size_t i = 0; i < idx.size(); ++i)
send_buffer[i] = x[idx[i]];
}
std::vector<std::int64_t> recv_buffer(sc.remote_indices_block().size());
MPI_Request request = MPI_REQUEST_NULL;
sc.scatter_fwd_begin(send_buffer.data(), recv_buffer.data(), 1, request);
// ... unrelated work can be done here while communication is
// in flight, but send_buffer/recv_buffer must not be touched ...
sc.scatter_fwd_end(request);
{
auto& idx = sc.remote_indices_block();
for (std::size_t i = 0; i < idx.size(); ++i)
x_ghost[idx[i]] = recv_buffer[i];
}

A reverse scatter follows the same pattern with the roles of ::local_indices_block and ::remote_indices_block, and of send_buffer and recv_buffer, swapped; see ::scatter_rev_begin and ::scatter_rev_end.

Template Parameters
ContainerContainer type for storing the 'local' and 'remote' indices. On CPUs this is normally std::vector<std::int32_t>. For GPUs the container should store the indices on the device, e.g. using thrust::device_vector<std::int32_t>.

Constructor & Destructor Documentation

◆ Scatterer() [1/4]

template<class Container = std::vector<std::int32_t>>
Scatterer ( const IndexMap & map)
inlineexplicit

Create a scatterer for data with a layout described by an IndexMap.

Note
Collective.
Parameters
[in]mapIndex map that describes the parallel layout of data.

◆ Scatterer() [2/4]

template<class Container = std::vector<std::int32_t>>
template<class U>
Scatterer ( const Scatterer< U > & s)
inline

Cast-copy constructor.

Create a copy of a Scatterer, where the copy uses a different storage container for indices that are used in MPI communication. Example usage includes creating from a CPU-suitable Scatterer a GPU-suitable Scatterer that can be used with GPU-aware MPI to move data between devices. This would be typical when copying a la::Vector or la::MatrixCSR to/from a GPU. When copying a vector or matrix to/from a GPU, the underlying Scatter that manages parallel communication will usually be copied too with a different storage container.

Note
Collective. The neighbourhood communicators are duplicated, so all ranks must make the copy together.
Parameters
sScatterer to copy

◆ Scatterer() [3/4]

template<class Container = std::vector<std::int32_t>>
Scatterer ( const Scatterer< Container > & scatterer)
default

Copy constructor

Note
Collective, as for the cast-copy constructor. Move instead where the original is no longer required.

◆ Scatterer() [4/4]

template<class Container = std::vector<std::int32_t>>
Scatterer ( Scatterer< Container > && scatterer)
default

Move constructor

Note
Not collective, unlike the copy constructors: the communicators are taken over rather than duplicated.

◆ ~Scatterer()

template<class Container = std::vector<std::int32_t>>
~Scatterer ( )
default

Destructor

Note
Collective, since the communicators are freed.

Member Function Documentation

◆ local_indices_block()

template<class Container = std::vector<std::int32_t>>
const container_type & local_indices_block ( ) const
inlinenoexcept

Array of indices for packing/unpacking owned data to/from a send/receive buffer.

For a forward scatter, the indices are used to copy required entries in the owned part of the data array into the appropriate position in a send buffer. For a reverse scatter, indices are used for assigning (accumulating) the receive buffer values into the correct position in the owned part of the data array.

The indices are in blocks, so for a block size bs a buffer holds bs values per index and must be bs * local_indices_block().size() long.

For a forward scatter, if x is the owned part of an array and send_buffer is the send buffer, send_buffer is packed such that:

auto& idx = scatterer.local_indices_block()
std::vector<T> send_buffer(bs * idx.size())
for (std::size_t i = 0; i < idx.size(); ++i)
    for (int j = 0; j < bs; ++j)
        send_buffer[i * bs + j] = x[idx[i] * bs + j];

For a reverse scatter, if recv_buffer is the received buffer, then x is updated by

auto& idx = scatterer.local_indices_block()
std::vector<T> recv_buffer(bs * idx.size())
for (std::size_t i = 0; i < idx.size(); ++i)
    for (int j = 0; j < bs; ++j)
        x[idx[i] * bs + j]
            = op(recv_buffer[i * bs + j], x[idx[i] * bs + j]);

where op is a binary operation, e.g. x[...] = buffer[...] or x[...] += buffer[...].

Returns
Indices container.

◆ operator=()

template<class Container = std::vector<std::int32_t>>
Scatterer & operator= ( Scatterer< Container > && scatterer)
default

Move assignment

Note
Collective if this Scatterer holds communicators, since assigning to it frees them.

◆ remote_indices_block()

template<class Container = std::vector<std::int32_t>>
const container_type & remote_indices_block ( ) const
inlinenoexcept

Array of indices for packing/unpacking ghost data to/from a send/receive buffer.

For a forward scatter, the indices are used to unpack received data into ghost entries. For a reverse scatter, indices are used to pack ghost entries into the send buffer.

The indices are in blocks, so for a block size bs a buffer holds bs values per index and must be bs * remote_indices_block().size() long.

For a forward scatter, if xg is the ghost part of the data array and recv_buffer is the receive buffer, xg is updated as

auto& idx = scatterer.remote_indices_block()
std::vector<T> recv_buffer(bs * idx.size())
for (std::size_t i = 0; i < idx.size(); ++i)
    for (int j = 0; j < bs; ++j)
        xg[idx[i] * bs + j] = recv_buffer[i * bs + j];

For a reverse scatter, if send_buffer is the send buffer, then send_buffer is packed such that:

auto& idx = scatterer.remote_indices_block()
std::vector<T> send_buffer(bs * idx.size())
for (std::size_t i = 0; i < idx.size(); ++i)
    for (int j = 0; j < bs; ++j)
        send_buffer[i * bs + j] = xg[idx[i] * bs + j];
Returns
Block indices container.

◆ scatter_fwd_begin()

template<class Container = std::vector<std::int32_t>>
template<typename T>
void scatter_fwd_begin ( const T * send_buffer,
T * recv_buffer,
int bs,
MPI_Request & request ) const
inline

Start a non-blocking neighbourhood collective exchange of owned data with the ranks that ghost it.

As ::scatter_fwd_begin_dtype, but with the MPI datatype built from a block size.

Parameters
[in]send_bufferPacked local data associated with each owned local index to be sent to processes where the data is ghosted. See Scatterer::local_indices_block for the order of the buffer and how to pack.
[in,out]recv_bufferBuffer for storing received data. See Scatterer::remote_indices_block for the order of the buffer and how to unpack.
[in]bsNumber of values per index map index (the block size). The buffers hold bs values for each index in ::local_indices_block and ::remote_indices_block respectively.
[out]requestHandle for tracking the status of the non-blocking communication. Any value passed in is overwritten, and MPI_REQUEST_NULL is returned when this rank has nothing to communicate. The same handle must be passed to Scatterer::scatter_fwd_end to complete the communication.

◆ scatter_fwd_begin_dtype()

template<class Container = std::vector<std::int32_t>>
template<typename T>
void scatter_fwd_begin_dtype ( const T * send_buffer,
T * recv_buffer,
MPI_Datatype type,
MPI_Request & request ) const
inline

Start a non-blocking neighbourhood collective exchange of owned data with the ranks that ghost it.

The communication is completed by calling Scatterer::scatter_fwd_end. See ::local_indices_block for how to pack send_buffer and ::remote_indices_block for how to unpack recv_buffer.

This is a differently named function rather than an overload of ::scatter_fwd_begin because the underlying type of MPI_Datatype is implementation-defined, and is an integer type in some MPI implementations, so overloading on int and MPI_Datatype is not portably unambiguous.

Note
Collective MPI operation. Every rank in the communicator must call this function, including ranks without neighbours.
The send and receive buffers must not be changed or accessed until after a call to Scatterer::scatter_fwd_end.
The pointers send_buffer and recv_buffer must be pointers to the data on the target device. E.g., if the send and receive buffers are allocated on a GPU, the send_buffer and recv_buffer should be device pointers.
Parameters
[in]send_bufferPacked local data associated with each owned local index to be sent to processes where the data is ghosted. See Scatterer::local_indices_block for the order of the buffer and how to pack.
[in,out]recv_bufferBuffer for storing received data. See Scatterer::remote_indices_block for the order of the buffer and how to unpack.
[in]typeMPI datatype for the data associated with one index, e.g. dolfinx::MPI::Datatype<T>(bs).type(). Buffer counts and displacements are in units of type, and the same type must be used on all ranks. MPI keeps a datatype alive until communication using it has completed, so type may be freed as soon as this function returns.
[out]requestHandle for tracking the status of the non-blocking communication. Any value passed in is overwritten, and MPI_REQUEST_NULL is returned when this rank has nothing to communicate. The same handle must be passed to Scatterer::scatter_fwd_end to complete the communication.

◆ scatter_fwd_end()

template<class Container = std::vector<std::int32_t>>
void scatter_fwd_end ( MPI_Request & request) const
inline

Complete a non-blocking MPI neighbourhood collective send.

This function completes the communication started by ::scatter_fwd_begin or ::scatter_fwd_begin_dtype.

Note
Local completion of the caller's own request, not itself collective. Every rank that called ::scatter_fwd_begin must still call this before reusing the buffers.
Parameters
[in,out]requestHandle returned by the matching begin call. Set to MPI_REQUEST_NULL once the communication has completed.

◆ scatter_rev_begin()

template<class Container = std::vector<std::int32_t>>
template<typename T>
void scatter_rev_begin ( const T * send_buffer,
T * recv_buffer,
int bs,
MPI_Request & request ) const
inline

Start a non-blocking neighbourhood collective exchange of ghost data with the owning ranks.

As ::scatter_rev_begin_dtype, but with the MPI datatype built from a block size.

Parameters
[in]send_bufferData associated with each ghost index. This data is sent to the process that owns the index. See Scatterer::remote_indices_block for the order of the buffer and how to pack.
[in,out]recv_bufferBuffer for storing received data. See Scatterer::local_indices_block for the order of the buffer and how to unpack.
[in]bsNumber of values per index map index (the block size). The buffers hold bs values for each index in ::remote_indices_block and ::local_indices_block respectively.
[out]requestHandle for tracking the status of the non-blocking communication. Any value passed in is overwritten, and MPI_REQUEST_NULL is returned when this rank has nothing to communicate. The same handle must be passed to Scatterer::scatter_rev_end to complete the communication.

◆ scatter_rev_begin_dtype()

template<class Container = std::vector<std::int32_t>>
template<typename T>
void scatter_rev_begin_dtype ( const T * send_buffer,
T * recv_buffer,
MPI_Datatype type,
MPI_Request & request ) const
inline

Start a non-blocking neighbourhood collective exchange of ghost data with the owning ranks.

The communication is completed by calling Scatterer::scatter_rev_end. See ::remote_indices_block for how to pack send_buffer and ::local_indices_block for how to unpack recv_buffer.

This is a differently named function rather than an overload of ::scatter_rev_begin because the underlying type of MPI_Datatype is implementation-defined, and is an integer type in some MPI implementations, so overloading on int and MPI_Datatype is not portably unambiguous.

Note
Collective MPI operation. Every rank in the communicator must call this function, including ranks without neighbours.
The send and receive buffers must not be changed or accessed until after a call to Scatterer::scatter_rev_end.
The pointers send_buffer and recv_buffer must be pointers to the data on the target device. E.g., if the send and receive buffers are allocated on a GPU, the send_buffer and recv_buffer should be device pointers.
Parameters
[in]send_bufferData associated with each ghost index. This data is sent to the process that owns the index. See Scatterer::remote_indices_block for the order of the buffer and how to pack.
[in,out]recv_bufferBuffer for storing received data. See Scatterer::local_indices_block for the order of the buffer and how to unpack.
[in]typeMPI datatype for the data associated with one index, e.g. dolfinx::MPI::Datatype<T>(bs).type(). Buffer counts and displacements are in units of type, and the same type must be used on all ranks. MPI keeps a datatype alive until communication using it has completed, so type may be freed as soon as this function returns.
[out]requestHandle for tracking the status of the non-blocking communication. Any value passed in is overwritten, and MPI_REQUEST_NULL is returned when this rank has nothing to communicate. The same handle must be passed to Scatterer::scatter_rev_end to complete the communication.

◆ scatter_rev_end()

template<class Container = std::vector<std::int32_t>>
void scatter_rev_end ( MPI_Request & request) const
inline

Complete a non-blocking MPI neighbourhood collective send.

This function completes the communication started by ::scatter_rev_begin or ::scatter_rev_begin_dtype.

Note
Local completion of the caller's own request, not itself collective. Every rank that called ::scatter_rev_begin must still call this before reusing the buffers.
Parameters
[in,out]requestHandle returned by the matching begin call. Set to MPI_REQUEST_NULL once the communication has completed.

The documentation for this class was generated from the following file: