shubhrapandit commited on
Commit
1d04d31
·
verified ·
1 Parent(s): 74536ef

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +16 -8
README.md CHANGED
@@ -162,11 +162,11 @@ The following performance benchmarks were conducted with [vLLM](https://docs.vll
162
  <th>Model</th>
163
  <th>Average Cost Reduction</th>
164
  <th>Latency (s)</th>
165
- <th>QPD</th>
166
- <th>Latency (s)th>
167
- <th>QPD</th>
168
  <th>Latency (s)</th>
169
- <th>QPD</th>
170
  </tr>
171
  </thead>
172
  <tbody style="text-align: center">
@@ -265,7 +265,9 @@ The following performance benchmarks were conducted with [vLLM](https://docs.vll
265
  </tbody>
266
  </table>
267
 
 
268
 
 
269
 
270
  ### Multi-stream asynchronous performance (measured with vLLM version 0.7.2)
271
 
@@ -284,11 +286,11 @@ The following performance benchmarks were conducted with [vLLM](https://docs.vll
284
  <th>Model</th>
285
  <th>Average Cost Reduction</th>
286
  <th>Maximum throughput (QPS)</th>
287
- <th>QPD</th>
288
  <th>Maximum throughput (QPS)</th>
289
- <th>QPD</th>
290
  <th>Maximum throughput (QPS)</th>
291
- <th>QPD</th>
292
  </tr>
293
  </thead>
294
  <tbody style="text-align: center">
@@ -386,4 +388,10 @@ The following performance benchmarks were conducted with [vLLM](https://docs.vll
386
  <td>4838</td>
387
  </tr>
388
  </tbody>
389
- </table>
 
 
 
 
 
 
 
162
  <th>Model</th>
163
  <th>Average Cost Reduction</th>
164
  <th>Latency (s)</th>
165
+ <th>Queries Per Dollar</th>
166
+ <th>Latency (s)<th>
167
+ <th>Queries Per Dollar</th>
168
  <th>Latency (s)</th>
169
+ <th>Queries Per Dollar</th>
170
  </tr>
171
  </thead>
172
  <tbody style="text-align: center">
 
265
  </tbody>
266
  </table>
267
 
268
+ **Use case profiles: Image Size (WxH) / prompt tokens / generation tokens
269
 
270
+ **QPD: Queries per dollar, based on on-demand cost at [Lambda Labs](https://lambdalabs.com/service/gpu-cloud) (observed on 2/18/2025).
271
 
272
  ### Multi-stream asynchronous performance (measured with vLLM version 0.7.2)
273
 
 
286
  <th>Model</th>
287
  <th>Average Cost Reduction</th>
288
  <th>Maximum throughput (QPS)</th>
289
+ <th>Queries Per Dollarv</th>
290
  <th>Maximum throughput (QPS)</th>
291
+ <th>Queries Per Dollar</th>
292
  <th>Maximum throughput (QPS)</th>
293
+ <th>Queries Per Dollar</th>
294
  </tr>
295
  </thead>
296
  <tbody style="text-align: center">
 
388
  <td>4838</td>
389
  </tr>
390
  </tbody>
391
+ </table>
392
+
393
+ **Use case profiles: Image Size (WxH) / prompt tokens / generation tokens
394
+
395
+ **QPS: Queries per second.
396
+
397
+ **QPD: Queries per dollar, based on on-demand cost at [Lambda Labs](https://lambdalabs.com/service/gpu-cloud) (observed on 2/18/2025).