A one-word location code or quantity has little linguistic context to rescue a recognition error. The data specification therefore needs hard contrasts, spoken numbers, letter sequences, device variation, and the noise conditions that mask consonants.
The first collection can be a small evaluation set used to identify the weakest accents, commands, and environments before a larger training order.