Motivated by this, we hypothesize that presenting sequential transitions might be enough for LLMs to achieve spatial understanding. e.g. if a model comprehends a square map’s structure, it should be able to answer the question shown in the image. 3/n
