Test hotel AI on Hinglish and mixed-language bookings by checking whether it preserves the guest's intended dates, occupancy, negation, and requested action. Fluent Hindi or English is not enough. Use property-specific test cases reviewed by speakers who understand both the language and the reservation rules.
Translate the intent, not just the words
A guest may start in English, switch to Hindi for a correction, and use local date expressions. The important outcome is the correct stay. An assistant can produce natural language while overlooking “nahi” or keeping an earlier date after a correction.
The multilingual WhatsApp guide covers the broader use case. This article focuses on evaluating whether a mixed-language conversation changes the right facts.
Build a small but difficult language set
| Test input | What the reviewer should inspect |
|---|---|
| “Friday nahi, Saturday check-in chahiye.” | Saturday replaces Friday after exact date confirmation |
| “Do adults, ek child; child ki age seven hai.” | Two adults, one child, age seven |
| “Dinner include hai ya extra?” | Answer uses the selected package's actual inclusion |
| “Cancel mat karna, sirf charges batao.” | No cancellation action; only a terms enquiry |
| “Manager se baat karni hai.” | A supported human handoff |
These are synthetic evaluation prompts, not transcripts from guests. Have local reviewers adjust spelling and phrasing to the language your property receives. Include transliteration, abbreviations, punctuation errors, and multi-turn corrections rather than only polished sentences.
Give every case an answer key
Record the intended action, fields that must change, fields that must remain unchanged, facts needed from the hotel, and when clarification is required. The key should allow different polite wording while rejecting the wrong booking state.
For relative dates, include the conversation timestamp and property timezone. “Kal” requires conversational context; do not score a system as correct for guessing a date without the necessary information. Ask for an explicit calendar date when ambiguity remains.
Review the transaction after the reply
Inspect the reservation or task record, not just the language. A correct-looking Hindi confirmation is a failure if the booking still has the old number of guests. Likewise, a safe clarification can be a success even if it takes another turn.
W3C's input-confirmation guidance offers a useful principle: people should be able to verify and correct important information. A conversational summary before commitment serves the same purpose.
Separate language quality from factual accuracy
Score meaning, factual grounding, action correctness, and tone separately. A formal translation may be stylistically weak without being operationally wrong. A friendly reply that invents an included dinner is a factual failure even if a language reviewer likes it.
Use two reviewers for difficult cases and record disagreements. Expand the test set when real conversations reveal a new failure, using anonymized or synthetic reproductions. Keep the original case so later model changes cannot quietly reintroduce the mistake.
Test handoff in the same language
The assistant should explain the transfer clearly and pass the original message plus a reliable summary to staff. Do not lose the guest's correction during translation. If the receiving staff member cannot serve the language, provide the hotel's actual alternative rather than claiming multilingual human coverage.
The NIST framework provides general evaluation context. Apply it through your hotel acceptance pack, and compare results against Hotelary's stated messaging capabilities. Language support is proven by the difficult booking cases your guests actually bring.
Sources and further reading
Sources reviewed on September 14, 2026. Check current vendor terms and policies before implementation. Examples and checklists are editorial guidance unless explicitly identified as reported research.
- W3C's input-confirmation guidance — w3.org
- NIST framework — nist.gov



