fix llama runner (#6256)
Summary:
Pull Request resolved: https://github.com/pytorch/executorch/pull/6256
imported-using-ghimport
Test Plan:
Imported from OSS
Run the following command and make sure it generates the right result:
```
python -m examples.models.llama2.runner.eager \
--checkpoint /home/lunwenh/models/1B/consolidated.00.pth \
--params /home/lunwenh/models/1B/params.json \
--max_len 128 \
--tokenizer /home/lunwenh/models/1B/tokenizer.model \
--prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>
You are a good assistant<|eot_id|><|start_header_id|>user<|end_header_id|>
What is the capital of France?<|eot_id|><|start_header_id|>assistant<|end_header_id|>"
```
```
Response:
The capital of France is Paris.
Tokens:
[791, 6864, 315, 9822, 374, 12366, 13]
```
Reviewed By: mergennachin
Differential Revision: D64442223
Pulled By: helunwencser
fbshipit-source-id: 99ee56f73e472a7243b8896a35ee092f287edf6b L
Lunwen He committed
1d7c3abf0c6d9b1c312e395b684e9c5bb87512a4
Parent: 2c8b14c
Committed by Facebook GitHub Bot <facebook-github-bot@users.noreply.github.com>
on 10/16/2024, 5:43:24 PM