MedSeek

Overview

MedSeek is a medical text retrieval encoder for representing clinical queries and reference cases. It produces embeddings for retrieving cases from clinical records. The workflow code is available in the MedSeek GitHub repository.

Released artifacts

This repository provides the MedSeek retrieval encoder weights, tokenizer, model configuration, and export metadata. Source code and sample data are available in the MedSeek GitHub repository.

Usage

Install the required packages:

python -m pip install torch transformers

Generate embeddings and calculate query-case similarity:

import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer

MODEL_ID = "Yjian1998/medseek"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModel.from_pretrained(MODEL_ID)
model.eval()


@torch.inference_mode()
def encode(texts, prefix):
    prepared = [f"{prefix}: {text}" for text in texts]
    inputs = tokenizer(
        prepared,
        max_length=512,
        padding=True,
        truncation=True,
        return_tensors="pt",
    )
    outputs = model(**inputs)
    mask = inputs["attention_mask"].unsqueeze(-1).bool()
    hidden = outputs.last_hidden_state.masked_fill(~mask, 0.0)
    pooled = hidden.sum(dim=1) / inputs["attention_mask"].sum(dim=1, keepdim=True)
    return F.normalize(pooled, p=2, dim=1)


query = ["Severe abdominal pain with hypotension and guarding."]
cases = [
    "Acute abdominal pain with shock and peritoneal signs.",
    "Mild intermittent abdominal discomfort with stable vital signs.",
]

query_embedding = encode(query, "query")
case_embeddings = encode(cases, "passage")
similarity = query_embedding @ case_embeddings.T
print(similarity)

Run the complete MedSeek retrieval example:

git clone https://github.com/YijianWu/medseek.git
cd medseek
python -m pip install -e .

python run_retrieval.py \
  --data data/data.csv \
  --model-path Yjian1998/medseek \
  --device auto \
  --output-dir outputs/demo

Intended use and limitations

MedSeek supports research on medical text representation, case retrieval, and similarity analysis. Performance can vary across institutions, languages, documentation practices, patient populations, and retrieval corpora.

Clinical use requires independent validation, appropriate data governance, regulatory review, and qualified professional oversight. Similarity scores should be interpreted together with the source records and the evaluation protocol used for each application.

Downloads last month
66
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Yjian1998/medseek

Finetuned
(192)
this model