Exploring LLMs and RAG at the 2025 AI Winter School (Brown University)

A workshop on querying physics papers with LLMs and checking the answers against retrieved source passages.

2025 AI Winter School banner — Brown University Department of Physics, Center for the Fundamental Physics of the Universe, January 13–16, 2025

At the 2025 AI Winter School, hosted by the Center for the Fundamental Physics of the Universe at Brown University, I led a 2.5-hour hands-on workshop on using large language models with physics-specific source material.

Participants compared direct model answers with answers generated from retrieved passages. We used LUX calibration papers and Brown Particle Astrophysics theses, asking about the D-D neutron energy, the electric fields used in yield measurements, and the energy and origin of low-energy 127Xe calibration events.

Checking numerical answers

For numerical answers, we checked the source passage, units, measurement conditions, and stated uncertainty. Page references made it possible to compare the answer with the original document.

Retrieval Model

The RAG system in the notebooks used a standard dense-retrieval pipeline:

  1. parse source documents from a Google Drive directory
  2. split the extracted text into overlapping chunks
  3. embed each chunk into a vector space
  4. embed the user question into the same vector space
  5. retrieve the top-ranked chunks by vector similarity
  6. pass those chunks, plus the question, to the LLM

For example, a dense retriever can rank a question q against a document chunk c_i using cosine similarity:

score(q, c_i) = cos(embed(q), embed(c_i))

Two Implementations

We provided two Colab notebooks: one used a hosted API, and the other ran an open-weight model in the Colab runtime.

Model Execution Retrieval
gpt-4o-mini OpenAI API LlamaIndex document loading, chunking, embeddings, vector index, and query engine
meta-llama/Meta-Llama-3.1-8B-Instruct Hugging Face model in a GPU-backed Colab runtime LlamaIndex with BAAI/bge-small-en-v1.5 embeddings

Indexing Parameters

The January 2025 notebooks pinned llama-index==0.12.3 and used the following settings:

Settings.chunk_size = 1000
Settings.chunk_overlap = 100

query_engine = index.as_query_engine(similarity_top_k=5)
response = query_engine.query(question)
  • Chunk size: Larger chunks preserve more surrounding text but use more of the prompt and make individual matches less specific.
  • Chunk overlap: Overlap retains context when a definition or explanation crosses a chunk boundary.
  • Embedding model: The embeddings determine which chunks are ranked near a question.
  • Top-k retrieval: Increasing similarity_top_k includes more candidate passages, at the cost of a longer prompt and potentially more irrelevant text.

PDF extraction may separate a table from its caption or detach units from the associated values. Chunk boundaries can compound the problem by splitting a definition from the passage that uses it.

Inspecting retrieved passages

The response.metadata and response.source_nodes fields identify the retrieved documents and passages:

response.metadata
response.source_nodes

The retrieved passages help distinguish three sources of error:

Source of error Description Check
Missing document The relevant paper or thesis is absent from the index. Loaded documents and index contents
Retrieval The relevant text is indexed, but the returned passages do not contain the answer. source_nodes, page metadata, chunk text, and ranking
Answer generation A retrieved passage contains the answer, but the model changes a value, unit, or qualification. The generated answer against the passage

Incremental Indexing

The notebooks also added new documents to an existing index. The first corpus used LUX D-D calibration papers; the second pass inserted Brown Particle Astrophysics theses:

new_documents = SimpleDirectoryReader(new_llama_index_data_path, recursive=True).load_data()
new_nodes = SimpleNodeParser().get_nodes_from_documents(new_documents)
index.insert_nodes(new_nodes)

Adding the missing thesis

The hosted-API notebook includes answers to the same question before and after the theses were added:

How low in energy was the ER response measured using 127Xe? Where did the 127Xe come from?

Before the theses were added, the retrieved D-D passages included a discussion of cosmogenic 131mXe. The model answered about that isotope and did not identify the 127Xe threshold.

After the theses were added, response.source_nodes included the relevant passages from the Huang thesis (pages 77–78 in the notebook metadata). The answer reported an energy deposition of 186 eV and attributed the 127Xe to cosmogenic activation while the xenon was above ground.

Materials