Mercurial
view asyncio_threads/inference/README.md @ 263:ee04e4e69fed
Add functional JRPG frame and tools
Add full-screen background_2 apertures, functional frame chrome, card-driven details, dual-window tools, live conversion workflows, and bounded cleanup for generated downloads.
Co-authored-by: Copilot <[email protected]>
| author | MrJuneJune <mrjunejune@users.noreply.github.com> |
|---|---|
| date | Thu, 06 Aug 2026 11:31:30 -0700 |
| parents | 46daba6e3cf4 |
| children |
line wrap: on
line source
Inference Questions Context You are tasked with building a simplified inference engine component responsible for handling incoming user requests for a large language model (LLM). To optimize throughput and GPU utilization, the engine must batch multiple requests together, run the inference call once per batch, and then deconstruct the results to return token-level output to the individual users. Objective Complete the provided Python class, BatchInferenceEngine by implementing the methods necessary to: Queue incoming user requests. Process a batch when the queue reaches a defined batch size. Simulate the token-level output from an LLM and correctly associate each generated token with its original request. Task Requirements Implement the logic for $enqueue\_request$. Implement the logic for $\_process\_batch$. Demonstrate the usage by creating 7 unique requests and enqueueing them one by one. Show the state of the queue and the processed tokens after each batch run.