The assistant reads MongoDB - and the data is still spreadsheets

FileToMongo » Use cases » Prepare data for a vector search app

Download FileToMongo Free Trial »     Buy FileToMongo Now »

The situation

An internal assistant, a support bot or a search feature answers out of a MongoDB collection: the application turns each document's text into an embedding, and MongoDB Atlas Vector Search (or the application's own index) finds the nearest ones. Everything from the embedding onward is the application team's job.

The part that keeps landing on a data person's desk is upstream. The material the assistant should know about lives where it was produced - a product catalogue as CSV, a policy list as Excel, a glossary or a parts list as JSON or XML - and it has to become documents in MongoDB with predictable field names before any of it can be embedded. It is also a job that repeats: the catalogue changes, the policies are revised, and the collection has to be brought up to date.

What you do

  1. Point FileToMongo at the file or folder - CSV, TSV, TXT, Excel, JSON, JSONL, XML, HTML, RDF or SQL. Field names come from the header row and the file encoding is detected, with a preview grid showing exactly what will be written.
  2. Map the columns to collection fields - keep the source names, or rename them to what the application expects. Choose the collection and load; millions of rows are streamed rather than held in memory.
  3. Load a whole folder in one run when the export produces one file per table, or append many files into a single collection.
  4. Re-run the same job when the exports refresh: save it as a session, or call it from the command line or a scheduled task.

Mapping spreadsheet columns to MongoDB collection fields in FileToMongo

The file is read locally and the documents are written to the MongoDB server you name - so the source data does not pass through anybody else's service on the way in.

Where this fits

FileToMongo stops at the collection. Turning a field into an embedding, choosing the chunk size, and building the vector index are the application's steps - Atlas Vector Search, a vector store, or an in-house index. This tool's job is to make the source data land as clean, complete, correctly named documents, so the indexing step has something to work with.

What it will not do

Related

All scenarios: FileToMongo use cases. The plain import: load a spreadsheet into MongoDB. Turning exports into JSON/JSONL first: DataFileConverter. Landing the data in a relational database instead - for pgvector: FileToDB. Repeating it: sessions, command line.