causal : llm is only able to access the token till the index i (i.e : till now ) . llm is forced to stop looking into the future tokens so it predicts it