<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>Paweł Dubiel Tech Blog</title>
  <link>https://paweldubiel.com</link>
  <atom:link href="https://paweldubiel.com/en/rss.xml" rel="self" type="application/rss+xml" />
  <description>A public digital garden about AI, software engineering, books, and learning.</description>
  <language>en</language>
  <lastBuildDate>Wed, 07 Oct 2026 00:00:00 GMT</lastBuildDate>
<item>
  <title>Sequence Packing</title>
  <link>https://paweldubiel.com/en/notes/42gq%E2%81%9D%20Sequence%20Packing</link>
  <guid>https://paweldubiel.com/en/notes/42gq%E2%81%9D%20Sequence%20Packing</guid>
  <pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate>
  <description>You are teaching a model to write short customer responses: “Thank you”, “No”, and “Yes”. After division into pieces of text, or tokens, the examples illustratively occupy 3, 2, and 2 positions, including each text&apos;s end</description>
</item>
<item>
  <title>Reinforcement Learning</title>
  <link>https://paweldubiel.com/en/notes/42gk%E2%81%9D%20Reinforcement%20Learning</link>
  <guid>https://paweldubiel.com/en/notes/42gk%E2%81%9D%20Reinforcement%20Learning</guid>
  <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
  <description>You want to teach a program to win a small game. At a fork, it can choose route A, collect two coins immediately, and finish the second turn without any new coins. Route B gives nothing on the first turn but leads to fiv</description>
</item>
<item>
  <title>Predicting command outputs helps with later agent training</title>
  <link>https://paweldubiel.com/en/notes/42eo%E2%81%9D%20Przewidywanie%20wynik%C3%B3w%20polece%C5%84%20pomaga%20w%20p%C3%B3%C5%BAniejszym%20treningu%20agenta</link>
  <guid>https://paweldubiel.com/en/notes/42eo%E2%81%9D%20Przewidywanie%20wynik%C3%B3w%20polece%C5%84%20pomaga%20w%20p%C3%B3%C5%BAniejszym%20treningu%20agenta</guid>
  <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
  <description>An agent working in a terminal tries to run a program. It sends a command, reads an error message and chooses its next action. In doing so, it must connect the output to the earlier command: recognize whether a file, mem</description>
</item>
<item>
  <title>Loss masking</title>
  <link>https://paweldubiel.com/en/notes/42eo1%E2%81%9D%20Loss%20masking</link>
  <guid>https://paweldubiel.com/en/notes/42eo1%E2%81%9D%20Loss%20masking</guid>
  <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
  <description>You are training an agent on records of conversations with tools. A record contains its commands and the tools&apos; responses. You want to assess selected parts while still letting it read the rest as context for its next ac</description>
</item>
<item>
  <title>pass@k</title>
  <link>https://paweldubiel.com/en/notes/42eo2%E2%81%9D%20pass%40k</link>
  <guid>https://paweldubiel.com/en/notes/42eo2%E2%81%9D%20pass%40k</guid>
  <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model produced ten programs, and one passed the tests. Does that mean a single response is likely to succeed? You need to distinguish “it worked on the first attempt” from “we found a success among many attempts”. pa</description>
</item>
<item>
  <title>Continuous Batching</title>
  <link>https://paweldubiel.com/en/notes/42el%E2%81%9D%20Continuous%20Batching</link>
  <guid>https://paweldubiel.com/en/notes/42el%E2%81%9D%20Continuous%20Batching</guid>
  <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
  <description>A server generates two responses at once. One finishes quickly; the other is long. If the next user must wait for the entire pair to finish, the freed slot is wasted for many steps. Continuous Batching allows the group o</description>
</item>
<item>
  <title>PagedAttention</title>
  <link>https://paweldubiel.com/en/notes/42ek%E2%81%9D%20PagedAttention</link>
  <guid>https://paweldubiel.com/en/notes/42ek%E2%81%9D%20PagedAttention</guid>
  <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
  <description>A server handles conversations of different lengths. If it reserves a large contiguous memory region in advance for each entire future response, some space remains unused. You want to allocate smaller pieces as needed. P</description>
</item>
<item>
  <title>Prefix Caching</title>
  <link>https://paweldubiel.com/en/notes/42em%E2%81%9D%20Prefix%20Caching</link>
  <guid>https://paweldubiel.com/en/notes/42em%E2%81%9D%20Prefix%20Caching</guid>
  <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
  <description>You ask five questions about the same long document. Each input starts with the same instruction and document text; only the ending contains a new question. Computing the same beginning again would repeat work. Prefix Ca</description>
</item>
<item>
  <title>Time to First Token</title>
  <link>https://paweldubiel.com/en/notes/42en%E2%81%9D%20Time%20to%20First%20Token</link>
  <guid>https://paweldubiel.com/en/notes/42en%E2%81%9D%20Time%20to%20First%20Token</guid>
  <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
  <description>You click “Send” and stare at an empty chat. For the user&apos;s comfort, it matters when the beginning of the response appears, even if writing the rest takes much longer. You need to measure that initial wait. Time to First</description>
</item>
<item>
  <title>Decode</title>
  <link>https://paweldubiel.com/en/notes/42eg%E2%81%9D%20Decode</link>
  <guid>https://paweldubiel.com/en/notes/42eg%E2%81%9D%20Decode</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>The first piece of the response has appeared. The next must account for what was just written: after “Today I drink tea” the model has a different context than after “Today I drink water”. This changing text must be exte</description>
</item>
<item>
  <title>FlashAttention</title>
  <link>https://paweldubiel.com/en/notes/42eh%E2%81%9D%20FlashAttention</link>
  <guid>https://paweldubiel.com/en/notes/42eh%E2%81%9D%20FlashAttention</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>Attention compares many pairs of text fragments. Saving all intermediate results to a computing card&apos;s large memory can make data transfers expensive. You want to perform the same calculation while reducing movement and </description>
</item>
<item>
  <title>Prefill</title>
  <link>https://paweldubiel.com/en/notes/42ef%E2%81%9D%20Prefill</link>
  <guid>https://paweldubiel.com/en/notes/42ef%E2%81%9D%20Prefill</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>You send a long question to a chat and wait for the first word. Before the model appends a response, it must process the content it has already received. This initial stage can require considerable work, especially for a</description>
</item>
<item>
  <title>Speculative Decoding</title>
  <link>https://paweldubiel.com/en/notes/42ei%E2%81%9D%20Speculative%20Decoding</link>
  <guid>https://paweldubiel.com/en/notes/42ei%E2%81%9D%20Speculative%20Decoding</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>A large model generates text correctly but slowly. A smaller model can quickly propose several following pieces. You want to use its work while keeping the large model in control of the final result. Speculative Decoding</description>
</item>
<item>
  <title>Beam Search</title>
  <link>https://paweldubiel.com/en/notes/42eb%E2%81%9D%20Beam%20Search</link>
  <guid>https://paweldubiel.com/en/notes/42eb%E2%81%9D%20Beam%20Search</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>When choosing the first piece of a response, you may reject too early a path that would later turn out better. Checking all possible sequences is too expensive, however. You need a limited pool of alternatives. Beam Sear</description>
</item>
<item>
  <title>Greedy Decoding</title>
  <link>https://paweldubiel.com/en/notes/42ea%E2%81%9D%20Greedy%20Decoding</link>
  <guid>https://paweldubiel.com/en/notes/42ea%E2%81%9D%20Greedy%20Decoding</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model could append several different pieces. The simplest choice is always to take the most likely one. This saves consideration of alternatives, but does the best first step lead to the best entire sequence? Greedy </description>
</item>
<item>
  <title>Grouped-Query Attention</title>
  <link>https://paweldubiel.com/en/notes/42ed%E2%81%9D%20Grouped-Query%20Attention</link>
  <guid>https://paweldubiel.com/en/notes/42ed%E2%81%9D%20Grouped-Query%20Attention</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model reads earlier text in several ways in parallel. Each such computation is an attention head. You want to reduce the memory these heads occupy, but sharing a single data set completely may constrain the model too</description>
</item>
<item>
  <title>Multi-Query Attention</title>
  <link>https://paweldubiel.com/en/notes/42ec%E2%81%9D%20Multi-Query%20Attention</link>
  <guid>https://paweldubiel.com/en/notes/42ec%E2%81%9D%20Multi-Query%20Attention</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model connects the current piece of text to the earlier conversation through several parallel computations. These computations are called attention heads. If each head stores its own descriptions of the entire preced</description>
</item>
<item>
  <title>Rotary Position Embedding</title>
  <link>https://paweldubiel.com/en/notes/42ee%E2%81%9D%20Rotary%20Position%20Embedding</link>
  <guid>https://paweldubiel.com/en/notes/42ee%E2%81%9D%20Rotary%20Position%20Embedding</guid>
  <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model compares pieces of text, but also needs to know how far apart they are. “Yesterday” beside a verb can play a different role from the same word in a distant quotation. How can position be included when comparing</description>
</item>
<item>
  <title>KL Divergence</title>
  <link>https://paweldubiel.com/en/notes/42dw%E2%81%9D%20KL%20Divergence</link>
  <guid>https://paweldubiel.com/en/notes/42dw%E2%81%9D%20KL%20Divergence</guid>
  <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
  <description>Two models divide probabilities differently among the same answers. One assigns 90% to A and 10% to B; the other gives each 50% . You want to measure how well the second reproduces the first model&apos;s predictions. KL Diver</description>
</item>
<item>
  <title>Knowledge Distillation</title>
  <link>https://paweldubiel.com/en/notes/42dx%E2%81%9D%20Knowledge%20Distillation</link>
  <guid>https://paweldubiel.com/en/notes/42dx%E2%81%9D%20Knowledge%20Distillation</guid>
  <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
  <description>A large model performs a task well but is too expensive for everyday use. A smaller model could try to reproduce its behavior. It needs examples or assessments showing more than just a list of correct labels, however. Kn</description>
</item>
<item>
  <title>KV Cache</title>
  <link>https://paweldubiel.com/en/notes/42dy%E2%81%9D%20KV%20Cache</link>
  <guid>https://paweldubiel.com/en/notes/42dy%E2%81%9D%20KV%20Cache</guid>
  <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model appends a response piece by piece. It has already computed much information about earlier text, so computing it from scratch at every step would repeat work. You want to retain results that remain valid. KV Cac</description>
</item>
<item>
  <title>Top-k Sampling</title>
  <link>https://paweldubiel.com/en/notes/42du%E2%81%9D%20Top-k%20Sampling</link>
  <guid>https://paweldubiel.com/en/notes/42du%E2%81%9D%20Top-k%20Sampling</guid>
  <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
  <description>When sampling the next piece, the model allows many extremely unlikely options. You want to keep some diversity but remove the long tail of candidates. You can admit only a specified number of leaders. Top k Sampling sam</description>
</item>
<item>
  <title>Top-p Sampling</title>
  <link>https://paweldubiel.com/en/notes/42dv%E2%81%9D%20Top-p%20Sampling</link>
  <guid>https://paweldubiel.com/en/notes/42dv%E2%81%9D%20Top-p%20Sampling</guid>
  <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
  <description>Sometimes the model has one clear leader; sometimes it has many continuations with similar scores. Always admitting five candidates ignores this difference. You want to choose the group&apos;s size according to where the prob</description>
</item>
<item>
  <title>Cross-entropy</title>
  <link>https://paweldubiel.com/en/notes/42do%E2%81%9D%20Cross-entropy</link>
  <guid>https://paweldubiel.com/en/notes/42do%E2%81%9D%20Cross-entropy</guid>
  <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
  <description>Two models identify the same word as the best continuation, but the first gives it 80% probability and the second only 20% . Checking only the winner misses this difference. Training needs a measure of how strongly the m</description>
</item>
<item>
  <title>LoRA</title>
  <link>https://paweldubiel.com/en/notes/42ds%E2%81%9D%20LoRA</link>
  <guid>https://paweldubiel.com/en/notes/42ds%E2%81%9D%20LoRA</guid>
  <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
  <description>You have a large model and want to adapt it to the style of your documents. Training all its weights — the numbers controlling its computations — requires a lot of memory. Can you train only a smaller correction to what </description>
</item>
<item>
  <title>Masked Language Modeling</title>
  <link>https://paweldubiel.com/en/notes/42dq%E2%81%9D%20Masked%20Language%20Modeling</link>
  <guid>https://paweldubiel.com/en/notes/42dq%E2%81%9D%20Masked%20Language%20Modeling</guid>
  <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
  <description>You want to teach the model to use the entire sentence. You hide a fragment in “Alice has a MASK ” or “Alice MASK a cat” and ask it to reconstruct the original. Information on both sides of the gap may be needed for the </description>
</item>
<item>
  <title>Perplexity</title>
  <link>https://paweldubiel.com/en/notes/42dp%E2%81%9D%20Perplexity</link>
  <guid>https://paweldubiel.com/en/notes/42dp%E2%81%9D%20Perplexity</guid>
  <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
  <description>You compare two models on the same text. You want to know which predicts its successive fragments better, rather than merely which guessed more individual winners. You need a number summarizing the probabilities assigned</description>
</item>
<item>
  <title>Temperature</title>
  <link>https://paweldubiel.com/en/notes/42dr%E2%81%9D%20Temperature</link>
  <guid>https://paweldubiel.com/en/notes/42dr%E2%81%9D%20Temperature</guid>
  <pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model usually chooses similar continuations, and you want to regulate sampling diversity. You do not change its knowledge or weights. You change only how strongly the best scoring candidates&apos; lead affects their proba</description>
</item>
<item>
  <title>Autoregressive Language Modeling</title>
  <link>https://paweldubiel.com/en/notes/42dj%E2%81%9D%20Autoregressive%20Language%20Modeling</link>
  <guid>https://paweldubiel.com/en/notes/42dj%E2%81%9D%20Autoregressive%20Language%20Modeling</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>You start a sentence: “Today I drink…”. After appending “tea”, possible continuations differ from those after “water”. The model must account for the text already chosen at every next step, rather than guess all position</description>
</item>
<item>
  <title>Byte Pair Encoding</title>
  <link>https://paweldubiel.com/en/notes/42dl%E2%81%9D%20Byte%20Pair%20Encoding</link>
  <guid>https://paweldubiel.com/en/notes/42dl%E2%81%9D%20Byte%20Pair%20Encoding</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>A vocabulary containing every possible word would be enormous and would continually lack new names. A vocabulary of individual characters is small but turns text into a very long sequence. You need a compromise: frequent</description>
</item>
<item>
  <title>Fine-tuning</title>
  <link>https://paweldubiel.com/en/notes/42dn%E2%81%9D%20Fine-tuning</link>
  <guid>https://paweldubiel.com/en/notes/42dn%E2%81%9D%20Fine-tuning</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>A trained model can process language, but you want it to classify company messages as “return”, “delivery”, or “payment”. General ability does not yet provide the way of working you expect. You prepare examples of messag</description>
</item>
<item>
  <title>Pre-training</title>
  <link>https://paweldubiel.com/en/notes/42dm%E2%81%9D%20Pre-training</link>
  <guid>https://paweldubiel.com/en/notes/42dm%E2%81%9D%20Pre-training</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>You want the model to recognize relationships in language before adapting it to a specific job, such as classifying customer messages. Manually preparing answers for all the text would be impractical. The text itself, ho</description>
</item>
<item>
  <title>Softmax</title>
  <link>https://paweldubiel.com/en/notes/42dk%E2%81%9D%20Softmax</link>
  <guid>https://paweldubiel.com/en/notes/42dk%E2%81%9D%20Softmax</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model scored three possible sentence continuations with the numbers 2 , 1 , and 0 . You want to sample a response according to their probabilities. Raw scores are not probabilities: they can be negative and do not su</description>
</item>
<item>
  <title>Causal Masking</title>
  <link>https://paweldubiel.com/en/notes/42dc%E2%81%9D%20Causal%20Masking</link>
  <guid>https://paweldubiel.com/en/notes/42dc%E2%81%9D%20Causal%20Masking</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>You are teaching the model to append “cat” after “Alice has a”. The word “cat” already exists in the complete example. If the model could look at it while predicting, the task would appear easy but would not resemble cre</description>
</item>
<item>
  <title>Cross-attention</title>
  <link>https://paweldubiel.com/en/notes/42dd%E2%81%9D%20Cross-attention</link>
  <guid>https://paweldubiel.com/en/notes/42dd%E2%81%9D%20Cross-attention</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>In translation, the model has two different texts: the original and the response still being created. When appending another translation word, it needs information from the original. Mixing only the response already writ</description>
</item>
<item>
  <title>Layer Normalization</title>
  <link>https://paweldubiel.com/en/notes/42df%E2%81%9D%20Layer%20Normalization</link>
  <guid>https://paweldubiel.com/en/notes/42df%E2%81%9D%20Layer%20Normalization</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>One fragment&apos;s description enters the next layer with numbers 2, 4 , another with 20, 40 . The signal magnitude can differ considerably. You want to aid computation by converting numbers within each description to a comp</description>
</item>
<item>
  <title>Position-wise Feed-Forward Network</title>
  <link>https://paweldubiel.com/en/notes/42de%E2%81%9D%20Position-wise%20Feed-Forward%20Network</link>
  <guid>https://paweldubiel.com/en/notes/42de%E2%81%9D%20Position-wise%20Feed-Forward%20Network</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model gathered information from the rest of the sentence for a particular fragment. It now needs to process that description: combine features and compute a new representation. Reading other positions again is not th</description>
</item>
<item>
  <title>Token Embedding</title>
  <link>https://paweldubiel.com/en/notes/42db%E2%81%9D%20Token%20Embedding</link>
  <guid>https://paweldubiel.com/en/notes/42db%E2%81%9D%20Token%20Embedding</guid>
  <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
  <description>The computer receives a number identifying a text fragment. The number 17 itself does not tell it what the fragment is or how it relates to others. It needs a useful numerical description that can be transformed during c</description>
</item>
<item>
  <title>Quantization errors partly cancel out across model layers</title>
  <link>https://paweldubiel.com/en/notes/42da%E2%81%9D%20B%C5%82%C4%99dy%20kwantyzacji%20cz%C4%99%C5%9Bciowo%20znosz%C4%85%20si%C4%99%20mi%C4%99dzy%20warstwami%20modelu</link>
  <guid>https://paweldubiel.com/en/notes/42da%E2%81%9D%20B%C5%82%C4%99dy%20kwantyzacji%20cz%C4%99%C5%9Bciowo%20znosz%C4%85%20si%C4%99%20mi%C4%99dzy%20warstwami%20modelu</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>You want to run a language model on your own graphics card, but its weights take up too much memory. One solution is post training quantization: you replace numbers stored in the finished model with values from a smaller</description>
</item>
<item>
  <title>Post-training quantization</title>
  <link>https://paweldubiel.com/en/notes/42da1%E2%81%9D%20Post-training%20quantization</link>
  <guid>https://paweldubiel.com/en/notes/42da1%E2%81%9D%20Post-training%20quantization</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>You already have a trained model, but it occupies too much memory to run on your chosen computer. Much of that space stores the numbers controlling computation. Can they be stored less precisely to shrink the model? Post</description>
</item>
<item>
  <title>Residual connection</title>
  <link>https://paweldubiel.com/en/notes/42da2%E2%81%9D%20Residual%20connection</link>
  <guid>https://paweldubiel.com/en/notes/42da2%E2%81%9D%20Residual%20connection</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model processes a text description through many layers. If every layer had to rebuild all the information, forwarding it could easily become difficult. Retaining the input and adding only the computed change is helpf</description>
</item>
<item>
  <title>LM head</title>
  <link>https://paweldubiel.com/en/notes/42da3%E2%81%9D%20LM%20head</link>
  <guid>https://paweldubiel.com/en/notes/42da3%E2%81%9D%20LM%20head</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model has processed “Today I drink” and has an internal description of this content as a list of numbers. The user, however, is waiting for text. We need a transition from the description hidden in the model to score</description>
</item>
<item>
  <title>Multi-Head Attention</title>
  <link>https://paweldubiel.com/en/notes/42cy%E2%81%9D%20Multi-Head%20Attention</link>
  <guid>https://paweldubiel.com/en/notes/42cy%E2%81%9D%20Multi-Head%20Attention</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>One combination of information can mix different sentence relationships too strongly. In translation, relationships involving people, time, and grammatical form may all be useful. We want to let the model compute several</description>
</item>
<item>
  <title>Positional Encoding</title>
  <link>https://paweldubiel.com/en/notes/42cz%E2%81%9D%20Positional%20Encoding</link>
  <guid>https://paweldubiel.com/en/notes/42cz%E2%81%9D%20Positional%20Encoding</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>You read “BMW overtook Volvo” and “Volvo overtook BMW”. They contain exactly the same words, but changing their order reverses the cars&apos; roles. If the model sees only word descriptions and compares them as an unordered s</description>
</item>
<item>
  <title>Scaled Dot-Product Attention</title>
  <link>https://paweldubiel.com/en/notes/42cx%E2%81%9D%20Scaled%20Dot-Product%20Attention</link>
  <guid>https://paweldubiel.com/en/notes/42cx%E2%81%9D%20Scaled%20Dot-Product%20Attention</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>The model needs to combine information from several sentence fragments. One is more useful, another less so. It needs a specific way to compute their contributions and then assemble the information into one result. Scale</description>
</item>
<item>
  <title>Self-attention</title>
  <link>https://paweldubiel.com/en/notes/42cw%E2%81%9D%20Self-attention</link>
  <guid>https://paweldubiel.com/en/notes/42cw%E2%81%9D%20Self-attention</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>In “Anna put the book down because it was heavy”, the word “it” alone explains little. It must connect to earlier fragments. A model processing text needs a way for one position&apos;s description to incorporate information f</description>
</item>
<item>
  <title>Transformer</title>
  <link>https://paweldubiel.com/en/notes/42cv%E2%81%9D%20Transformer</link>
  <guid>https://paweldubiel.com/en/notes/42cv%E2%81%9D%20Transformer</guid>
  <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
  <description>You want to translate “The train will arrive tomorrow”. Translating each word separately is insufficient: the correct form depends on other words and their order. The model needs a way to combine information from text, e</description>
</item>
<item>
  <title>The same ranking yields different decisions after reordering</title>
  <link>https://paweldubiel.com/en/notes/42be%E2%81%9D%20Ten%20sam%20ranking%20daje%20inne%20decyzje%20po%20zmianie%20kolejno%C5%9Bci</link>
  <guid>https://paweldubiel.com/en/notes/42be%E2%81%9D%20Ten%20sam%20ranking%20daje%20inne%20decyzje%20po%20zmianie%20kolejno%C5%9Bci</guid>
  <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
  <description>A system answering questions from documents often operates in four steps. A search engine finds candidates, a scorer assigns each one a number indicating its usefulness, a threshold rejects low scores, and a reader model</description>
</item>
<item>
  <title>Batched pointwise scoring</title>
  <link>https://paweldubiel.com/en/notes/42be1%E2%81%9D%20Batched%20pointwise%20scoring</link>
  <guid>https://paweldubiel.com/en/notes/42be1%E2%81%9D%20Batched%20pointwise%20scoring</guid>
  <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
  <description>You have a question and three retrieved documents. You want to score each one&apos;s usefulness, but three separate model calls repeat the same instruction and question. You can show all three together and ask for three score</description>
</item>
<item>
  <title>nDCG@10</title>
  <link>https://paweldubiel.com/en/notes/42be2%E2%81%9D%20nDCG%20at%2010</link>
  <guid>https://paweldubiel.com/en/notes/42be2%E2%81%9D%20nDCG%20at%2010</guid>
  <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
  <description>A search engine returns ten results. Good documents may be on the list, but if the most useful is only tenth, the reader may never reach it. You need a measure accounting for both relevance and position. nDCG@10 measures</description>
</item>
</channel>
</rss>