Skip to content

Benchmark ​

This page explains how to run the benchmark in benchmark/movielens/ and records one historical result. The numbers apply only to the specified hardware, data, and deployment configuration; they are not a performance guarantee for other environments.

What the Benchmark Covers ​

The benchmark calls the main_rec API and includes:

  • four recall paths: global popularity, user genres, ItemCF, and Milvus vector search;
  • exposure deduplication and recall merging;
  • item-detail lookup and genre diversification;
  • asynchronous recommendation-log writes to Kafka;
  • exposure-history writes to Redis.

init.sh runs init_model.sql to train, export, and deploy the recall and ranking models. Loading item features calls the DSSM item tower Service and writes its embeddings to Milvus. These preparation steps are outside the timed wrk run. By default, benchmark requests do not call an online model Service.

Dataset and Flow ​

The benchmark uses MovieLens 1M:

ItemSize
Users6,040
Items3,706 movies
RatingsAbout 1 million
Embedding dimensions64

The recommendation flow is defined in benchmark/movielens/init_sqlrec_sql.sql.

StageDefault behavior
Global popularityRead 300 rows from global_hot_item
Genre recallQuery genre_hot_item from user_interest_genre, limit 300
ItemCF recallQuery itemcf_i2i from user_recent_click_item, limit 300
Vector recallQuery Milvus with a local random user vector, limit 300
DeduplicationRemove items exposed within the last hour
Rankingrank_fun_simple joins item information without inference
DiversificationWindow 3, at most 1 item per genre, return 10

request.lua explicitly sets use_recall_service=false and rank_fun=rank_fun_simple to keep the timed workload stable. To exercise online inference, set request params to {"use_recall_service":"true","rank_fun":"rank_fun"}. That path calls the DSSM user tower and Wide & Deep ranking Services. Random user vectors in the default benchmark are for performance measurement, not recommendation-quality evaluation.

Prerequisites ​

Start the complete environment by following Service Deployment. init.sh also uses or installs wrk, starts Kyuubi, and downloads MovieLens data and Python dependencies.

The historical result below used:

ItemConfiguration
CPUAMD Ryzen 5600H
Memory32 GB DDR4
Operating systemDebian 12
DeploymentMinikube, one SQLRec instance

Initialize the Data ​

From the repository root:

bash
cd benchmark/movielens
bash init.sh

The initialization script and mvn test share the repository root .venv. The script creates it when needed and installs benchmark dependencies. Set PYTHON_BOOTSTRAP to choose the base Python interpreter (3.10 or newer).

The script:

  1. Deploys Kyuubi and prepares wrk.
  2. Creates the Milvus item_embedding collection and index.
  3. Downloads MovieLens 1M, converts it to Parquet, and uploads it to HDFS.
  4. Creates offline and online connector tables.
  5. Computes popularity, genre, and ItemCF features with Spark SQL.
  6. Trains, exports, and deploys the Wide & Deep and DSSM models.
  7. Writes features to Redis and model-generated item vectors to Milvus.
  8. Registers SQL functions and the main_rec API.
  9. Calls main_rec through Beeline in both default and model-backed modes as basic checks.

init.sh changes HDFS, Redis, Milvus, Kafka, and SQLRec metadata in the target environment. Do not run it against a shared or production environment.

Run the Benchmark ​

bash
cd benchmark/movielens
bash benchmark.sh

The current script uses a 10-second warm-up with one thread and connection, followed by a 30-second run with 10 threads and 10 connections against /api/v1/main_rec. Each request chooses a valid MovieLens user ID from 1 to 6040.

Treat the current benchmark.sh and request.lua as authoritative. Record concurrency and duration with the result whenever you change them.

Historical Result ​

The following result was measured with the earlier mock-item-vector setup and a different user-ID range. The current setup imports model-generated item vectors, so the numbers are not directly comparable.

text
Running 30s test @ http://192.168.49.2:30001/api/v1/main_rec
  10 threads and 10 connections
  Thread Stats   Avg      Stdev     Max   +/- Stdev
    Latency     6.73ms    3.16ms  90.29ms   94.46%
    Req/Sec   151.20     16.58   191.00     73.67%
  45231 requests in 30.02s, 87.90MB read
Requests/sec:   1506.47
Transfer/sec:      2.93MB
MetricValueMeaning
Average latency6.73 mswrk request latency
Latency standard deviation3.16 mswrk statistic
Maximum latency90.29 msMaximum observed in this run
Per-thread requests/sec151.20Average in Thread Stats, not total QPS
Total requests45,231During 30.02 seconds
Total QPS1,506.47Requests/sec
Transfer rate2.93 MB/sTransfer/sec

When comparing runs, keep at least the dataset, SQLRec version, JVM, concurrency, connector deployment, and model-service usage consistent.