![]() |
|
#2
|
|||
|
|||
|
I mean, this is precisely how a model trained to predict the next toen with attention layers and thinking modes would be expected to work when the training data is full of such strong patterns. I would be more surprised knowing the transformer architecture that it did not work this way.
|
|
|