[WIP] FLAN-T5 integration #194

afeldman-nm · 2024-04-17T19:16:41Z

FILL IN THE PR DESCRIPTION HERE

FIX #xxxx (link existing issues this PR will resolve)

BEFORE SUBMITTING, PLEASE READ THE CHECKLIST BELOW AND FILL IN THE DESCRIPTION ABOVE

PR Checklist (Click to Expand)

Thank you for your contribution to vLLM! Before submitting the pull request, please ensure the PR meets the following criteria. This helps vLLM maintain the code quality and improve the efficiency of the review process.

PR Title and Classification

Only specific types of PRs will be reviewed. The PR title is prefixed appropriately to indicate the type of change. Please use one of the following:

[Bugfix] for bug fixes.
[CI/Build] for build or continuous integration improvements.
[Doc] for documentation fixes and improvements.
[Model] for adding a new model or improving an existing model. Model name should appear in the title.
[Frontend] For changes on the vLLM frontend (e.g., OpenAI API server, LLM class, etc.)
[Kernel] for changes affecting CUDA kernels or other compute kernels.
[Core] for changes in the core vLLM logic (e.g., LLMEngine, AsyncLLMEngine, Scheduler, etc.)
[Hardware][Vendor] for hardware-specific changes. Vendor name should appear in the prefix (e.g., [Hardware][AMD]).
[Misc] for PRs that do not fit the above categories. Please use this sparingly.

Note: If the PR spans more than one category, please include all relevant prefixes.

Code Quality

The PR need to meet the following code quality standards:

We adhere to Google Python style guide and Google C++ style guide.
Pass all linter checks. Please use format.sh to format your code.
The code need to be well-documented to ensure future contributors can easily understand the code.
Include sufficient tests to ensure the project to stay correct and robust. This includes both unit tests and integration tests.
Please add documentation to docs/source/ if the PR modifies the user-facing behaviors of vLLM. It helps vLLM user understand and utilize the new features or changes.

Notes for Large Changes

Please keep the changes as concise as possible. For major architectural changes (>500 LOC excluding kernel/data/config/test), we would expect a GitHub issue (RFC) discussing the technical design and justification. Otherwise, we will tag it with rfc-required and might not go through the PR.

What to Expect for the Reviews

The goal of the vLLM team is to be a transparent reviewing machine. We would like to make the review process transparent and efficient and make sure no contributor feel confused or frustrated. However, the vLLM team is small, so we need to prioritize some PRs over others. Here is what you can expect from the review process:

After the PR is submitted, the PR will be assigned to a reviewer. Every reviewer will pick up the PRs based on their expertise and availability.
After the PR is assigned, the reviewer will provide status update every 2-3 days. If the PR is not reviewed within 7 days, please feel free to ping the reviewer or the vLLM team.
After the review, the reviewer will put an action-required label on the PR if there are changes required. The contributor should address the comments and ping the reviewer to re-review the PR.
Please respond to all comments within a reasonable time frame. If a comment isn't clear or you disagree with a suggestion, feel free to ask for clarification or discuss the suggestion.

Thank You

Finally, thank you for taking the time to read these guidelines and for your interest in contributing to vLLM. Your contributions make vLLM a great tool for everyone!

T5 enc/dec example file; linting/formatting

Small PR for debug print statements

…l_runner.py

fix _make_tensor_with_pad args change which broke decoder scenarios

…nce constructor call takes is_encoder_decoder, eos_token_id, lora_request calls; set is_encoder_decoder field in constructor

…token_id, lora_request arguments

… with import of PagedAttentionImpl

… prefix caching; low confidence of success

…py to model_runner.py

…ecutor/layers/attention

…s are still incorrect

…ables arguments to override input_metadata values; tests still pass but enc/dec still fails

…etc.

…ompt_lens is treated as a list in T5

…oder mode; removed encoder/decoder argument of Sequence

…on of relative position encoding based on packed-variable-length-sequences

…ct T5 inference result. Nothing is broken by this commit, unless there is a subsequent commit with changes in order to pass regression tests.

…ks wrong though. Added not_causal option for attn_bias to kernel interface contracts; also switched to batch size 1 to avoid incorrectness likely caused by packed-variable-sequence-length mask having zeroes rather than -inf's

…adata has correct blocktable, slot_mapping=None, and correct (max) context length(s) (derived from prompt); decode-phase decoder self-attention relative position encoding mask has 1 x K geometry where 1 is the number of new tokens generated in a step and K is context length padded to the nearest multiple of block size, and also mask is reshuffled with contiguous (); ensured general correctness of cross-attention input_metadata; modified T5 example script to prevent HF/vLLM T5 instances from being length limited; net effect: batch-size 1 seems to work but batch-size >1 not supported

Jin Shang and others added 30 commits February 29, 2024 09:27

t5-small

dd82ba3

fix

f2fd579

lint

2fb6905

T5 enc/dec example file; linting/formatting

be58c3b

native/vllm t5 comparison test

70837fd

merged upstream-main into enc_dec_t5

42a6e2b

Merge branch 'upstream-main' into enc_dec_t5

e3fd30d

Merge pull request #1 from afeldman-nm/enc_dec_t5

db726e6

T5 enc/dec example file; linting/formatting

remove debug print statements

43e920e

silence warning; legacy=False for tokenizer; lint/format

431f014

Merge branch 'js8544_enc_dec_t5' into enc_dec_t5

37fcf99

Merge pull request #2 from afeldman-nm/enc_dec_t5

4bf056b

Small PR for debug print statements

fix _make_tensor_with_pad args change which broke decoder scenario

8a5060f

fixed bug caused by non-handling of self.model_config is None in mode…

29d6f44

…l_runner.py

remove commented-out print statements

a4950ba

small cleanup

9c03760

Merge pull request #3 from afeldman-nm/enc_dec_t5

9f20ccf

fix _make_tensor_with_pad args change which broke decoder scenarios

arg naming fix

6d6dccd

Merge branch 'js8544_enc_dec_t5' into enc_dec_t5

7035178

fixed attention_kernels.cu merge conflict; questions about ROCM

dbec357

llm_engine.py conflict resolution; removed prefix caching code; Seque…

4b2a121

…nce constructor call takes is_encoder_decoder, eos_token_id, lora_request calls; set is_encoder_decoder field in constructor

actually updated Sequence constructor to take i_encoder_decoder, eos_…

a93c17d

…token_id, lora_request arguments

xformers.py accept incoming changes; replace paged_attention function…

a62c3af

… with import of PagedAttentionImpl

saved changed to xformers woops

c31921f

attempt at fixing model_runner conflicts related to encoder/decoder &…

0c78be9

… prefix caching; low confidence of success

encoder/decoder + prefix caching not supported; moved check from llm.…

e25e6b8

…py to model_runner.py

refactoring, including: moved enc_dec_attention.py into vllm/model_ex…

7f70d76

…ecutor/layers/attention

existing regressions pass (yay) but encoder/decoder example fails

36c8291

fixed encoder/decoder reshape and cache bug, but paged attention call…

08f268a

…s are still incorrect

augmented paged attention with context_lens, max_context_len, block_t…

b9b0600

…ables arguments to override input_metadata values; tests still pass but enc/dec still fails

afeldman-nm added 30 commits March 22, 2024 14:35

scheduler schedule() support cross block-tables and cross sequences, …

691c2c1

…etc.

LLMEngine can build a sequencegroup with cross sequences

e240eb4

t5 Sampler does not pass vocab size to constructor; input_metadata.pr…

cbfba8e

…ompt_lens is treated as a list in T5

add_request now correctly swaps decoder_prompt, prompt in encoder/dec…

501551c

…oder mode; removed encoder/decoder argument of Sequence

Added cross_input_metadata field to InputMetadata

08435e4

wip multi blocktable

6e459a2

wip

8e1ca33

plumbing dummy input metadata structures into model

e097732

plumbed encoder/decoder input metadata all the way into t5

2a44585

first pass at T5 encoder support

91a4608

inefficient but effective & Attention-wrapper-compatible implementati…

d0c5e36

…on of relative position encoding based on packed-variable-length-sequences

wip cross-attention

3737d5b

first pass at enc/dec support that runs e2e but doesn't produce corre…

38946ed

…ct T5 inference result. Nothing is broken by this commit, unless there is a subsequent commit with changes in order to pass regression tests.

to pass regression tests: removed debug prints

3c39f55

wip vllm, examples => fp32

4ec2fde

works on bsz = 1

38f55ed

wip

c1258b4

passing with t5-small

0af1022

refactoring out print statements

f5242a0

fix to pass regression tests

de0fd31

WIP google/flan-t5-xxxx

5a67647

removed print statement

ed05d47

batched enc/dec example

d5a8b92

wip, trying prompt padding

f555f5d

bs >1 prefill works

2c12b44

small change to examples

dba02b2

fix to support case where num prompts != 2

db201b6

set up (failing) flan-t5 test

ead7c82

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[WIP] FLAN-T5 integration #194

[WIP] FLAN-T5 integration #194

afeldman-nm commented Apr 17, 2024

[WIP] FLAN-T5 integration #194

Are you sure you want to change the base?

[WIP] FLAN-T5 integration #194

Conversation

afeldman-nm commented Apr 17, 2024

PR Title and Classification

Code Quality

Notes for Large Changes

What to Expect for the Reviews

Thank You