BSGS Algorithm Implementation for Nvidia GPUs

20 replies 78 views
dan88Member
Posts: 10 · Reputation: 40
#1Jun 8, 2018, 08:41 PM
Just put together my version of the BigStepGiantStep algorithm for Nvidia cards. Works on Cuda and Windows x64. Check it out here and share your speed results.
7 Reply Quote Share
ryan1337Newbie
Posts: 2 · Reputation: 18
#2Jun 9, 2018, 08:06 PM
Looks like you're familiar with BSGS and x86 assembler... I modified Jean's BSGS for curve 'r1' while btc uses k1. What’s the deal with the start value? Does the searched k need to be in that range?
1 Reply Quote Share
dan.walletSenior Member
Posts: 9 · Reputation: 848
#3Jun 9, 2018, 10:26 PM
Sweet, I’ll update you on speeds across different GPUs once I run some tests.
6 Reply Quote Share
stacksatsHero Member
Posts: 168 · Reputation: 2023
#4Jun 10, 2018, 04:30 AM
For sure. Your project seems way more user-friendly than that JLP kangaroo thing. Tuning that one is such a hassle!
1 Reply Quote Share
Posts: 1 · Reputation: 109
#5Jun 10, 2018, 05:37 AM
Just tried out your BSGS on my GTX 1660s and it was slower than JeanLucPons' Kangaroo. BSGS-cuda hit 330 Mkey/s, while Kangaroo got 450 Mkey/s.
3 Reply Quote Share
stacksatsHero Member
Posts: 168 · Reputation: 2023
#6Jun 10, 2018, 03:17 PM
We really need to see how long it takes to find a sample private key. What’s the fastest code for this?
0 Reply Quote Share
cyberbitMember
Posts: 76 · Reputation: 68
#7Jun 10, 2018, 07:40 PM
COBRAS, how about you jump in and start testing? Or are you waiting for someone else to do it?
2 Reply Quote Share
dan88Member
Posts: 10 · Reputation: 40
#8Jun 11, 2018, 01:03 AM
With v1.2 on a single 2080ti, I solved some pubkeys in about 28 minutes with params -w 26. Compared to JLP's CPU version, that’s six times faster!
5 Reply Quote Share
ryan1337Newbie
Posts: 2 · Reputation: 18
#9Jun 11, 2018, 06:02 AM
Nice! But how quick would it run with the full range from zeros to f's?
4 Reply Quote Share
jakechadSenior Member
Posts: 1 · Reputation: 1992
#10Jun 13, 2018, 05:42 PM
How much memory do we need for each babystep? Is the hashtable using GPU memory or global RAM?
2 Reply Quote Share
cyberbitMember
Posts: 76 · Reputation: 68
#11Jun 13, 2018, 09:36 PM
Lol, definitely some sarcasm there. COBRAS doesn’t seem to contribute much. His last comment made no sense at all.
4 Reply Quote Share
dan.walletSenior Member
Posts: 9 · Reputation: 848
#12Jun 14, 2018, 01:51 AM
Had my RTX 3070 running at 1,000 MKey/s with default settings. Didn’t mess with anything to optimize.
0 Reply Quote Share
dan.walletSenior Member
Posts: 9 · Reputation: 848
#13Jun 14, 2018, 02:26 AM
I ran the same benchmarks as Etar and JLP for 16 pubkeys. JLP on CPU took 3 hours and 35 minutes for comparison.
6 Reply Quote Share
dan88Member
Posts: 10 · Reputation: 40
#14Jun 14, 2018, 08:07 AM
Each baby step uses 8 bytes. The hashtable is stored in the GPU memory. With -w 26 and -htsz 25, it generates quite a bit of babysteps.
6 Reply Quote Share
Posts: 3 · Reputation: 20
#15Jun 14, 2018, 11:55 AM
Thanks, Etar! I think your BSGS-cuda outperforms JLP's version. JLP is solid but takes way too long on my GPU.
2 Reply Quote Share
dan88Member
Posts: 10 · Reputation: 40
#16Jun 14, 2018, 02:12 PM
Mandatory update v1.2.1 is out now. Fixed some bugs with multi-GPU searching.
2 Reply Quote Share
dan.walletSenior Member
Posts: 9 · Reputation: 848
#17Jun 14, 2018, 08:29 PM
JLP's BSGS is CPU only, while yours supports GPU. Side by side tests show that BSGS Cuda is faster for multiple pubkeys.
4 Reply Quote Share
Posts: 3 · Reputation: 20
#18Jun 15, 2018, 01:26 AM
You’re right. Sorry for the confusion. I just meant it was slow on my laptop, so slow that I sometimes just quit waiting.
3 Reply Quote Share
Posts: 3 · Reputation: 20
#19Jun 16, 2018, 07:24 AM
Works well in the 65-bit range, but hitting 120-bit seems limited. Still can’t get a key in the 65-bit range.
3 Reply Quote Share
LoneChadFull Member
Posts: 1 · Reputation: 478
#20Jun 16, 2018, 10:40 PM
Searching these two pubkeys in the 100-bit range, anyone got hardware info for this? What GPU models are you using for those results?
4 Reply Quote Share

Related topics