CodingBox Q&A Ask question

MCX516A-CCAT mesh 100G DAC पर: lshw 40Gbit/s बताता है और एक iperf3 stream 21 Gbit/s पर tops out करता है

Asked Active Viewed 49 AI translation from English
2

हमारे पास तीन-node cluster है, nodes आपस में सीधे 100G DAC से cabled हैं, बीच में कोई switch नहीं, और interfaces को broadcast bond में डाला है ताकि replication traffic का अपना fabric हो। असली load डालने से पहले मुझे एक baseline चाहिए था, और numbers आपस में match नहीं कर रहे।

  • 3x Mellanox ConnectX-5 EN, MCX516A-CCAT, dual-port QSFP28
  • हर node pair के बीच एक 100G DAC
  • AMD EPYC hosts, तीनों पर Proxmox

ethtool पूरी तरह खुश है:

# ethtool ens1
Settings for ens1:
        Supported link modes:   100000baseCR4/Full
        Advertised link modes:  100000baseCR4/Full
        Speed: 100000Mb/s
        Duplex: Full
        Link detected: yes

lshw नहीं है:

# lshw -class network
  *-network
       description: Ethernet interface
       vendor: Mellanox Technologies
       capacity: 40Gbit/s

और दो nodes के बीच iperf3 करीब 21 Gbit/s पर बैठता है, जो इन दोनों figures के आसपास भी नहीं है।

अब तक क्या किया:

  • दोनों cards पर cable को दूसरे port पर move किया, कोई फर्क नहीं
  • same type का दूसरा DAC लगाकर देखा, कोई फर्क नहीं
  • link पूरे समय up रहता है, counters पर कोई error नहीं चढ़ता

तो दोनों में से कौन सा tool मुझसे झूठ बोल रहा है, और मुझे cable, card या driver में से किसके पीछे जाना चाहिए?

Comments 5

Accepted answer

जो तुमने post किया उसमें कुछ भी DAC की तरफ इशारा नहीं करता। यहां दो अलग-अलग, एक-दूसरे से unrelated चीज़ें हो रही हैं।

पहली, disagreement। lshw -class network एक capability figure print करता है जो वह खुद निकालता है, और इन cards पर यह 100G पर आए link के बारे में खुशी-खुशी capacity: 40Gbit/s कह देगा। negotiated rate वो है जो ethtool report करता है, और तुम्हारा Speed: 100000Mb/s कहता है, 100000baseCR4/Full advertised के साथ। इस तरफ सब ठीक है, यहां कुछ fix करने को नहीं है।

दूसरी, throughput। कुछ और छूने से पहले slot check करो:

# lspci -vv
        LnkCap: ... Speed 8GT/s ...
        LnkSta: ... Speed 2.5GT/s ... (downgraded)

अगर LnkSta 2.5GT/s पर trained हुआ जबकि LnkCap कहता है 8GT/s, तो तुम wire से काफी नीचे capped हो और कितना भी cable बदलने से मदद नहीं मिलेगी। card reseat करो और confirm करो कि वह ऐसे slot में बैठा है जो असल में पूरी width के लिए wired है।

फिर single stream से measure करना बंद करो:

# iperf3 -P 8 -c <peer>

करीब 21 Gbit/s लगभग वही है जो इस class के host पर एक core तुम्हें देगा, तो अकेले वह number ज्यादा कुछ नहीं बताता। run के दौरान CPU देखो और यह भी देखो कि idle states क्या कर रहे हैं: bursts के बीच deep C-states में गिरते cores इस rate पर असली bandwidth खा जाते हैं।

3 Taiwanlinkeng56TW Show original (English) AI translation

कुछ भी replacement order करने से पहले, उस card के लिए lspci -vv से LnkSta line post करो और वह exact iperf3 command line भी जो इस्तेमाल की। 100G पर एक stream एक ही CPU core measure करता है, link नहीं, और लोग इस पर दिन बर्बाद कर देते हैं। और confirm करो कि test के दौरान दूसरा port वाकई mesh की दूसरी leg carry कर रहा है, idle बैठा नहीं है: एक slot से दो live 100G ports feed करना एक port से अलग budget है। मैं फिलहाल lshw को parked रखूंगा, यह इस सवाल के लिए सही tool नहीं है।

3 United Arab Emirateslambdahawk88AE Show original (English) AI translation

C-state वाले हिस्से पर एक correction: अगर hosts EPYC हैं, तो हर thread में copy-paste होने वाले intel_idle knobs तुम्हारे लिए कुछ नहीं करते, वो driver AMD पर path में है ही नहीं। जो lever मेरे लिए काम आया वो kernel command line पर processor.max_cstate=2 था। Same idea, अलग platform। बाकी उस post का सही है, खासकर यह कि negotiated rate lshw से मत पढ़ो।

2 SpainoptictechES Show original (English) AI translation

थोड़ी अलग failure, hardware का same family, slot सुलझने के बाद यह rule out करने लायक है: ConnectX-5 QSFP28 ports का direct mesh layer 3 पर बहुत आसानी से गलत हो जाता है। मेरे पास तीन nodes थे MCX516A-CCA_Ax पर, firmware 16.35.4030 DOCA 2.8.0 driver के साथ, MCP1600-C003E30L 3 m copper DAC से wired। हर link 100 Gbps पर active report करता था और एक भी ping पार नहीं हुई। सभी छह mesh interfaces के addresses एक ही 10.5.5.x subnet से थे, बीच में कहीं कोई switch नहीं, तो kernel के पास यह तय करने का कोई तरीका नहीं था कि कोई destination किस physical port का है। हर node pair के लिए एक subnet, 10.5.5.x, 10.5.6.x और 10.5.7.x, और यह काम करने लगा। किसी के copper को दोष देने से पहले तीनों boxes पर ip a और ip route चलाना बनता है।

1 Ukrainerxnode71UA Show original (English) AI translation

दोनों बातें सही निकलीं। lspci -vv ने दिखाया कि card LnkCap के 8GT/s के मुकाबले 2.5GT/s पर trained था, तो वही suspect number one था। तीनों boxes पर cards को दूसरे slots में move किया, LnkSta अब 8GT/s पर आता है, और iperf3 -P 8 के साथ वही pair of nodes तुरंत single-stream वाले figure से आगे निकल गया।

मैंने broadcast bond भी छोड़ दिया और mesh को RSTP के साथ Open vSwitch पर फिर से बनाया। तीन CPU threads और दोनों ports पर फैले iperf के साथ अब मैं करीब 95 Gbit/s measure कर रहा हूं, जो इस cluster के काम के लिए line rate के काफी करीब है। lshw अब भी 40Gbit/s पर अड़ा है और मैंने उसे देखना बंद कर दिया है।

1 United Kingdomcoaxpilot98GB Show original (English) AI translation
Log in to comment. Log in