SIGN IN SIGN UP

update eager runner to use same options for loading the model (#6257)

Summary:
Pull Request resolved: https://github.com/pytorch/executorch/pull/6257

imported-using-ghimport

Test Plan:
Imported from OSS

Run the following command and make sure it generate the right result:
```
python -m examples.models.llama2.runner.eager \
          -c /home/lunwenh/models/1B_Instruct/consolidated.00.pth \
          -p /home/lunwenh/models/1B_Instruct/params.json \
          -t /home/lunwenh/models/1B_Instruct/tokenizer.model \
          --max_seq_length 128 \
          -kv \
          --prompt "<|begin_of_text|><|start_header_id|>system<|end_header_id|>

You are a good assistant<|eot_id|><|start_header_id|>user<|end_header_id|>

What is the capital of France?<|eot_id|><|start_header_id|>assistant<|end_header_id|>"
```
```
Response:
The capital of France is Paris.<|eot_id|>
 Tokens:
 [791, 6864, 315, 9822, 374, 12366, 16134, 91, 68, 354, 851, 91, 29]
```

Reviewed By: mergennachin

Differential Revision: D64442224

Pulled By: helunwencser

fbshipit-source-id: bb8b11de6325ae76423b086491094a4444249553
L
Lunwen He committed
423f65d742d0914ed8cc147cc7f7fd92764a4758
Parent: 1d7c3ab
Committed by Facebook GitHub Bot <facebook-github-bot@users.noreply.github.com> on 10/16/2024, 5:43:24 PM