Skip to main content

lib.native_transform.torch.id_hash_tokenizer

ID-hash tokenizer layer for native feature transforms.

Maps arbitrary integer input values to contiguous, zero-based indices based on a provided vocabulary, mapping out-of-vocabulary values to a dedicated unknown index. The layer is TorchScript- and ONNX-exportable so it can be embedded in a model graph and run identically at train and serve time.

IDHashTokenizer Objects​

class IDHashTokenizer(nn.Module)

Map integer IDs to contiguous vocabulary indices.

TODO(https://github.com/michelangelo-ai/michelangelo/issues/1699): this core is a reusable, general-purpose layer colocated here only because native_transform is currently its sole consumer. Move it into a standalone reusable-layers package if a second consumer needs it or when another such layer is added.

Maps arbitrary input integer values to new, contiguous integer indices based on a provided vocabulary. Values not found in the vocabulary are mapped to an unknown index, which is set to the size of the (deduplicated) vocabulary.

The input vocabulary may be unsorted. The mapping from an original vocabulary value to its new index is based on its position in the provided vocabulary list (i.e. vocabulary[i] maps to i). Internally the values are sorted for an efficient torch.bucketize lookup, then remapped back to their original positions, so ordering of the provided list is preserved in the output indices.

The layer is compatible with both TorchScript and ONNX export.

Despite the name "Hash", this performs an exact vocabulary lookup via torch.bucketize (not a hash); the name is retained for backward compatibility.

Arguments:

  • vocabulary - List of integer values to map to contiguous indices. Duplicate values are removed, preserving the index of their first occurrence.

Raises:

  • TypeError - If vocabulary is not a list of integers.
  • ValueError - If vocabulary is empty.

Example:

>>> tokenizer = IDHashTokenizer(vocabulary=[-10, -3, 0, 2, 4, 6])
>>> tokenizer(torch.tensor([-10, 0, 5], dtype=torch.long))
tensor([0, 2, 6])

__init__​

def __init__(vocabulary: list[int]) -> None

Initialize the tokenizer from a vocabulary of integer values.

Arguments:

  • vocabulary - List of integer values to map to contiguous indices. Duplicate values are removed, preserving the index of their first occurrence.

Raises:

  • TypeError - If vocabulary is not a list of integers.
  • ValueError - If vocabulary is empty.

forward​

def forward(input_ids: torch.Tensor) -> torch.Tensor

Map input integer IDs to contiguous vocabulary indices.

Values not found in the vocabulary are mapped to unk_index.

Arguments:

  • input_ids - Tensor of integer IDs of any shape (e.g. (batch_size, sequence_length)). Must have dtype torch.int32 or torch.long.

Returns:

Tensor of mapped indices with the same shape and dtype as input_ids.

Raises:

  • TypeError - If input_ids is not of integer type (torch.int32 or torch.long).