Hello! I noticed that the model reads the chat_template during evaluation. For models like GLM-5.2 that use <|im_end|> as an end-of-sequence (EOS) token, mentioning chat_template during its thinking/reasoning phase will leads it to generate <|im_end|>, causing generation to terminate prematurely.
How do you evaluate models that use <|im_end|> as an EOS token, such as Kimi and GLM?
Hello! I noticed that the model reads the chat_template during evaluation. For models like GLM-5.2 that use <|im_end|> as an end-of-sequence (EOS) token, mentioning chat_template during its thinking/reasoning phase will leads it to generate <|im_end|>, causing generation to terminate prematurely.
How do you evaluate models that use <|im_end|> as an EOS token, such as Kimi and GLM?