The passage needs a home
Splitting a document creates smaller units for retrieval, but a passage can become misleading if it loses its title, section, date, or relationship to surrounding material. Preserve enough source identity to interpret the text after retrieval.
At minimum, decide how the application will identify the original document and the selected passage. Use stable identifiers that survive a display-title change or a different ordering of results.
Choose boundaries around meaning
A fixed size can be a useful processing constraint, but it does not determine where a sentence, exception, or procedure begins and ends. Inspect a sample of the resulting passages before indexing the full corpus.
Look for detached conditions, headings without their body, and instructions split away from the thing they apply to. The right preparation depends on the structure of your source material rather than one universal chunk size.
Preserve versions and relationships
When a document changes, the application needs to know which passages belong to which version and what should happen to superseded material. Keep that lifecycle explicit in the source model.
Useful metadata can also support grouping. If several relevant passages come from the same document, the application may want grouped retrieval to avoid crowding out other sources that contribute distinct evidence.
Test citations as part of retrieval
Take a returned passage and follow its identifier back to the source. Check that a reader can find the actual supporting text and understand the surrounding section.
Then test a question requiring an exception or a reference to another document. If the source identity and relationships are preserved, the application has a better foundation for retrieving and presenting the additional context.
Explore the next step
Continue with the interactive workflow. For the current setup and API contract, use the Polygres documentation.