lib.native_transform.torch.id_hash_tokenizer
ID-hash tokenizer layer for native feature transforms.
Maps arbitrary integer input values to contiguous, zero-based indices based on a provided vocabulary, mapping out-of-vocabulary values to a dedicated unknown index. The layer is TorchScript- and ONNX-exportable so it can be embedded in a model graph and run identically at train and serve time.
IDHashTokenizer Objects
class IDHashTokenizer(nn.Module)
Map integer IDs to contiguous vocabulary indices.
TODO(https://github.com/michelangelo-ai/michelangelo/issues/1699): this
core is a reusable, general-purpose layer colocated here only because
native_transform is currently its sole consumer. Move it into a
standalone reusable-layers package if a second consumer needs it or when
another such layer is added.
Maps arbitrary input integer values to new, contiguous integer indices based on a provided vocabulary. Values not found in the vocabulary are mapped to an unknown index, which is set to the size of the (deduplicated) vocabulary.
The input vocabulary may be unsorted. The mapping from an original
vocabulary value to its new index is based on its position in the provided
vocabulary list (i.e. vocabulary[i] maps to i). Internally the values
are sorted for an efficient torch.bucketize lookup, then remapped back
to their original positions, so ordering of the provided list is preserved in
the output indices.
The layer is compatible with both TorchScript and ONNX export.
Despite the name "Hash", this performs an exact vocabulary lookup via
torch.bucketize (not a hash); the name is retained for backward
compatibility.
Arguments:
vocabulary- List of integer values to map to contiguous indices. Duplicate values are removed, preserving the index of their first occurrence.
Raises:
TypeError- Ifvocabularyis not a list of integers.ValueError- Ifvocabularyis empty.
Example:
>>> tokenizer = IDHashTokenizer(vocabulary=[-10, -3, 0, 2, 4, 6])
>>> tokenizer(torch.tensor([-10, 0, 5], dtype=torch.long))
tensor([0, 2, 6])
__init__
def __init__(vocabulary: list[int]) -> None
Initialize the tokenizer from a vocabulary of integer values.
Arguments:
vocabulary- List of integer values to map to contiguous indices. Duplicate values are removed, preserving the index of their first occurrence.
Raises:
TypeError- Ifvocabularyis not a list of integers.ValueError- Ifvocabularyis empty.
forward
def forward(input_ids: torch.Tensor) -> torch.Tensor
Map input integer IDs to contiguous vocabulary indices.
Values not found in the vocabulary are mapped to unk_index.
Arguments:
input_ids- Tensor of integer IDs of any shape (e.g.(batch_size, sequence_length)). Must have dtypetorch.int32ortorch.long.
Returns:
Tensor of mapped indices with the same shape and dtype as
input_ids.
Raises:
TypeError- Ifinput_idsis not of integer type (torch.int32ortorch.long).