# Hdock3 scoring for large-scale complex models

**URL:** <https://ask.bioexcel.eu/t/hdock3-scoring-for-large-scale-complex-models/5592>\
**Category:** HADDOCK\
**Tags:** haddock\
**Created:** [April 11, 2025, 5:06am UTC](https://ask.bioexcel.eu/t/hdock3-scoring-for-large-scale-complex-models/5592 "2025-04-11T05:06:27Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dongl](https://avatars.discourse-cdn.com/v4/letter/d/7ea924/32.png) [@dongl](https://ask.bioexcel.eu/u/dongl)\
**Post date:** [April 11, 2025, 5:06am UTC](https://ask.bioexcel.eu/t/hdock3-scoring-for-large-scale-complex-models/5592/1 "2025-04-11T05:06:27Z")

</div>

I would like to perform large-scale scoring of my protein models using HADDOCK3. I have approximately 10,000 models that I need to evaluate. What would be the best strategy or recommended approach to handle this scale efficiently

---

<div class="post-metadata">

**Author:** ![VGPReys](https://dub1.discourse-cdn.com/flex013/user_avatar/ask.bioexcel.eu/vgpreys/32/960_2.png) [@VGPReys](https://ask.bioexcel.eu/u/VGPReys)\
**Post date:** [April 11, 2025, 6:46am UTC](https://ask.bioexcel.eu/t/hdock3-scoring-for-large-scale-complex-models/5592/2 "2025-04-11T06:46:53Z")

</div>

Dear dongl,

Haddock3 is limited to 20 input files, but unlimited in the number of conformations in an input ensemble.

To process 10000 models, I would suggest to merge them together in a single file using `pdb_mkensemble`.  
Then, you can process them using a standard scoring workflow:

```toml
run_dir = "big_scoring_run"
molecules = "10000_ensemble.pdb"

# Generation of topologies
[topoaa]

# A energy minimisation step followed by a scoring using the HADDOCK scoring function
[emscoring]

# Clustering by Fraction of common contacts
[clustfcc]
clust_cutoff = 0.9 # Group together models having >= 90% similar contacts
min_population = 1 # This parameter allows to keep all models even the ones that are not clustered (singlotons)

# Grouping models by clusters
[seletopclusts]
top_cluster = 10000 # in case they are all different
top_models = 10000 # in case they are all fall in the same cluster

# A final analysis step to generate the plots
[caprieval]

```

For reference, please see what we did for the scoring challenge in CAPRI rounds using haddock3 [10.1002/prot.26789](https://onlinelibrary.wiley.com/doi/10.1002/prot.26789).
