『Shopify's "Gisting": Squeezing a 6,000-Token Prompt Down to 1,500』のカバーアート

Shopify's "Gisting": Squeezing a 6,000-Token Prompt Down to 1,500

Shopify's "Gisting": Squeezing a 6,000-Token Prompt Down to 1,500

無料で聴く

ポッドキャストの詳細を見る

Jordan and Riley break down Shopify Engineering's "gisting" technique — replacing a big, repeated system prompt with a handful of learned "gist tokens" that carry the same behavior to the model, trained via self-distillation (teacher pass with the full prompt, student pass with the gist tokens, minimize the gap between them). They cover why deployment stays trivial — the learned embeddings are written straight into the model's embedding matrix, no custom serving path — and the payoff on Shopify's Sidekick GraphQL agent: 4:1 compression, time-to-first-token down 19%, end-to-end latency down 38%, and 14% fewer GPUs. This is original commentary and discussion, not a reproduction of the original post.

Source: "Gisting: Compressing LLM Agent Context to increase throughput and reduce cost" by Paige Vegna and Cody Mazza-Anthony, Shopify Engineering, Aug 19 2026 — https://shopify.engineering/gisting

adbl_web_anon_alc_button_suppression_t1
まだレビューはありません