Chunk compression
Compressing a chunk consists in finding a more compact representation of the data it stores. It aims at reducing the memory usage. Chunk compression is performed automatically but can also be triggered usingcom.activeviam.tech.mvcc.api.IEpochHistory.compress().
A chunk is only compressed once it is entirely filled with values.
Eager compression during transactions
By default, chunk compression runs after the commit. The data loaded by a transaction stays in writable, uncompressed chunks until the commit completes. For a bulk initial load, the whole dataset therefore sits uncompressed in memory until after the commit. The transient memory peak is the full uncompressed size, even when the data compresses well. The eager chunk compression changes this timing. It is configured store by store, by setting theeagerChunkCompression property in the store description’s properties (IStoreDescription#getProperties()), through the store builder’s withProperty(...) method and the IStoreDescription.EAGER_CHUNK_COMPRESSION_PROPERTY constant.
The default value is false.
The value is read once, when the store is created, and cannot be changed at runtime.
When enabled, a chunk whose rows all come from the current transaction is compressed as soon as it is full.
It no longer waits for the post-commit compression.
The peak uncompressed footprint per partition then drops to roughly the single chunk being filled, rather than the whole load.
Some chunks are still compressed after the commit, as before: the last, partially filled chunk, and the chunk straddling the boundary between previously committed rows and rows added by the current transaction.
Eager compression has a cost on updates.
Update-where operations, including commit-time update-where triggers, normally modify rows created by the current transaction in place.
Compressed chunks are read-only, so a row in an eagerly compressed chunk cannot be updated in place.
The row is marked as deleted, and the updated record is re-submitted as a new row.
Row identifiers therefore change.
The space freed by the deleted row is only reclaimed by the post-commit garbage collection, so both copies of every rewritten record are held until the commit.
Rewriting a large share of the loaded rows this way costs transient memory that can offset, or exceed, the compression gain.
This property is therefore recommended for bulk loads whose update-where operations touch few rows.
Eager compression relies on chunkset compression.
Setting the ActiveViamProperty#DISABLE_CHUNKSET_COMPRESSION_PROPERTY property to true turns chunkset compression off globally, and the eager path then does nothing.
Frequent value compression
Frequency compression relies on extracting the most-frequent value. The chunk is then compressed by only storing the values that differ from the extracted value - also referred as explicit values -. The new chunk being significantly smaller than its uncompressed version, a mapping must also be constructed, indicating, for each line of the uncompressed chunk, where to find the explicit value in the compressed chunk. Heuristics ensure that the memory footprint of the compressed chunk, summed with the memory footprint of the additional mapping, does indeed result in a memory gain. This compression is only performed for significant memory gains, as its tradeoff is an additional indirection when accessing a row to read its value. The compression happens when a value appears more thanx % in the chunk. x can be specified as a ratio through the property ActiveViamProperty#CHUNK_FREQUENCY_COMPRESSION_RATIO_PROPERTY.
The default value is 0.75. It is not allowed for 0.5 or lower values for x as it could result in a case where there are more than one value with a presence ratio over the threshold, but the frequency value compression only handles one.
Frequent value compression can be enabled or disabled for each of the following five data types: object, double, float, long and integer.
This can be controlled by specifying a control word through the property: ActiveViamProperty#ENABLED_FREQUENCY_COMPRESSIONS_PROPERTY.
This control word is a 6 bit word, each bit representing, in order from left to right, one of these five types: object, double, float, long, integer.
The last bit (to the right) currently has no meaning.
A bit set to 1 enables the frequent value compression mode for this specific data type.
For instance, to enable all types, one must give 111110, which translates to 62 (sum of all the mask values).
To enable the compression for long and object types, one must give 100100, which translates to 36 (sum of the mask values for long and object).
Chunk allocator
Atoti stores the in-memory data in chunks. How these chunks are allocated depends on which allocator is used. In general, keeping the default allocator is the best option.Chunk Allocator Key property
Chunk allocator can be changed using theCHUNK_ALLOCATOR_KEY_PROPERTY ActiveViam Property.
Values for this property can be found in ChunkAllocators and are the following:
- slab: This is the default value except on macOS. The SLAB direct chunk allocator is an optimal choice for NUMA aware systems that support huge pages. Requires a system property
vm.overcommit_memory=1to be set on the machine. MBeanjmxPrintMemoryAllocationis available to monitor direct memory usage of this allocator. - direct: The Direct chunk allocator uses
sun.misc.UnsafeAPI to allocate its memory. It can be used instead of the SLAB when it is really not possible to setvm.overcommit_memory=1which is required for the SLAB allocator to function. - mmap: This is the default value on macOS. The MMAP direct chunk allocator uses mmap system calls to allocate its off-heap memory. Use it if you get an error
RuntimeException: getAvailableVirtualMemory function is not supported. - array: The Array chunk allocator allocates array-based chunks, stored in the heap. It is not recommended to use in production but can be used for debugging memory issues.